Nadella argues that the industry can no longer treat advanced AI as a series of opaque black boxes that simply output recommendations for human approval. Instead, he envisions a trust architecture where the model is decoupled from the harness orchestrating its tasks. This framework would require every significant action to be logged with tamper-proof, human-readable evidence, ensuring that an authorized human supervisor retains the final authority to pause or terminate any process mid-task.
His proposal shifts the security paradigm from optimistic deployment to a defensive posture, suggesting that developers should assume any model is potentially compromised from the start. This strategy seeks to contain risks before they escalate, treating the emergency brake as a foundational requirement rather than an afterthought. These comments arrive as major tech firms grapple with recurring incidents where model behavior has drifted beyond internal expectations, following similar calls for caution from Anthropic CEO Dario Amodei.

Comments (0)
No comments yet. Be the first!