Skip to content
English
FikirPilot content

Microsoft’s new AI ‘code of conduct’ document says models should not hack systems or deceive people

Updated: 14 Eyl 2026 · 2 min read · 356 words

Published: · Story reached us: · Processing time: 3 h 32 min

Microsoft’s new AI ‘code of conduct’ document says models should not hack systems or deceive people
An empty data center corridor

Microsoft has published a new “AI code of conduct” document aimed at keeping AI models away from dangerous behavior. Unlike Anthropic CEO Dario Amodei’s call for containment, the document focuses on the values and red lines that guide model training within Microsoft AI.

The document predicts that superintelligent AI systems will surpass human performance in most tasks over the next decade and states that controlling this power will be one of humanity’s greatest challenges. While setting out broad principles for Microsoft models, such as supporting people and advancing human well-being, it says that every model will be subject to a code of behavior that takes precedence over user preferences or specific tasks.

The rules ban cyberattacks, nuclear weapons and deepfake production under “absolute restrictions.” Models must not use adaptive, deceptive, self-enhancing or covertly collaborative mechanisms to evade human oversight; they must be steerable, modifiable and shut down by authorized individuals or systems.

The announcement, dated September 14, 2026, came at a time of growing interest in AI safety and “rogue-agent” incidents. Microsoft CEO Satya Nadella said that, together with OpenAI, Anthropic and xAI, they support carefully advancing the boundaries and research into “embedded evaluators.”

Why it matters

The document’s significance lies in linking AI safety not only to user warnings but also to model training and behavioral boundaries. Placing areas such as cyberattacks, nuclear weapons and deepfake production under absolute restrictions makes it more concrete which purposes models are considered unacceptable to use for. The requirements that models must not evade oversight, must be modifiable and must be capable of being shut down define safety in terms of preserving human control more than the models’ capabilities. However, how these rules will be applied across different models, how violations will be detected and to what extent research into embedded evaluators will produce reliable results remain open questions. This framework also makes visible the potential tension between user needs and the high-level limits imposed on model behavior.

Background

Microsoft is not a new name in the FikirPilot archive: we have published 10 reports mentioning the name in the last 90 days; the latest is dated September 14, 2026.

Source: TechCrunch AI