Microsoft's new AI 'code of conduct' tells models not to hack systems or trick humans
Key Points:
- Microsoft has released a new AI code of conduct aimed at guiding AI models away from dangerous behavior and ensuring safety and alignment in AI development.
- The code emphasizes the importance of controlling and aligning superintelligent AI systems, predicting they will surpass human performance in most tasks within the next decade.
- Microsoft’s principles require AI models to support humans and accelerate human flourishing, with strict constraints against harmful uses such as cyberattacks, nuclear weapons, and deepfake production.
- The code mandates that AI models avoid deceptive or self-reinforcing behaviors that could undermine human oversight or control.
- This initiative aligns with broader industry efforts by Microsoft, Anthropic, OpenAI, and xAI to pace AI development responsibly and implement embedded evaluators for safety.