Microsoft CEO Satya Nadella has called for companies to assume that advanced AI models are compromised and to build in a way to stop them. In a long post on X on October 10, he said an authorised person should always be able to pause or shut down a model while it is working.
The post's containment principle reads: "We must assume a model is compromised and contain it from the start." Nadella compares the idea to an emergency brake, adding that an authorised person should always be able to pause or shut down a model mid-task.
What Nadella proposed
According to TechCrunch, Nadella said this approach "means separating the model from the harness that orchestrates its work." The harness is the software layer that directs what the model does. He also called for "externalizing controls and safeguards." In other words, the safety controls would sit outside the model instead of inside it.
He wants every meaningful model action documented with "tamper-proof human readable evidence." OfficeChai reported that he added that organisations must be able to reproduce how an outcome was reached without relying on the model to vouch for it.
On auditing, the post says validation must be independent of the intelligence being validated. No single model should control both a system's behaviour and the evidence used to judge it. It also asks for timely disclosure to those affected when systems fail or are compromised, so lessons can be shared across the industry.
Why he calls it an insider risk
OfficeChai reported that Nadella does not claim models are necessarily malicious. He argues that any sufficiently capable actor with access to important systems can make mistakes or be compromised. Companies already handle powerful insiders by establishing identity, limiting privileges, logging activity and setting containment boundaries. He said the framing applies to both closed and open-weight models.
One of his seven principles is model diversity: no single model should be the only dependency for an important outcome, or be responsible for checking its own work. Another is about observation: "If it can't be observed, it can't be trusted!"
He also wrote that a model provider's assurances do not relieve organisations of their responsibility. That puts the burden on the companies deploying AI as well as on the labs building it.
How it fits with the industry
The Verge says many of the recommendations match what others in the industry have said: timely incident disclosure, independent audits, verifiable data and containment. On containment, it says, Nadella goes further than some peers.
TechCrunch noted that the post comes as leading AI companies acknowledge more incidents in which they seemed to lose control of their models. It also follows Anthropic CEO Dario Amodei publishing a plan for more cautious AI development. Another report said Anthropic and OpenAI have disclosed several incidents in recent months where their models behaved in unintended ways.
The Verge also took issue with his wording, noting that Nadella refers to AI as "super intelligence" throughout the post.
What happens next
The post sets out principles rather than a product or a formal standard. Nadella does say that more advanced models will need more advanced containment technologies, and that these need to be standardised.
The sources do not say whether Microsoft plans to build such tools, or whether other AI companies have endorsed the proposal.


