Microsoft’s AI Code of Conduct Sets Safety Red Lines for Models

Microsoft has published a new AI code of conduct that sets out safety red lines for its AI models, moving from broad principles to specific technical constraints. The document, released amid growing industry concern over rogue agents and existential risk, is framed around the prediction that superintelligent AI will surpass humans in most tasks within the next decade. Microsoft argues that containing such systems is one of humanity’s greatest challenges, so it is clarifying both the purpose of building them and how they will be controlled.

The code of conduct defines an overarching framework for every Microsoft AI model, which takes precedence over individual user preferences or task-specific instructions. It includes “absolute constraints” that forbid cyberattacks, nuclear weapons, and deepfake production. Beyond these hard prohibitions, the document adds broader provisions to prevent loss of human control. Specifically, models are barred from using adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight, so they cannot become impossible to direct, modify, or shut down by authorized people or systems.

Microsoft‘s approach aligns with a general industry trend toward pacing frontier development, alongside Anthropic, OpenAI, and xAI. CEO Satya Nadella welcomed research and deliberate pacing aimed at getting alignment right, and expressed support for the concept of “embedded evaluators” as a mechanism to make safety commitments concrete. The code of conduct is positioned as a practical, training-level guide, complementing more philosophical calls for caution. It is intended to translate high-level values—like supporting humans rather than replacing them—into enforceable training constraints for Microsoft‘s models.

Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans | TechCrunch

View Original