09/15/2026 | Press release | Distributed by Public on 09/15/2026 05:41
Microsoft has published a provisional code of conduct for its artificial intelligence models, setting out restrictions on how its systems should behave as concerns grow over the risks posed by sophisticated AI models.
The guidelines come days after leaders at Anthropic and OpenAI backed the idea of deliberately slowing the pace of frontier AI development, and as researchers and policymakers intensify calls for stronger safeguards around advanced models.
For Microsoft, the move also provides an opportunity to define how it wants its own AI systems to operate as the company plays two roles in the industry: building its own models while serving as a major cloud provider and commercial partner to leading AI laboratories.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
Register for Nigeria Capital Market Masterclass.
Mustafa Suleyman, who leads Microsoft's model development, said the company had been working on the guidelines for roughly five months but decided to publish them now because of the recent debate over AI safety.
"We got feedback from people that they wanted to see even more explicit commitment to AI always working in service of people and not trying to replace them," Suleyman told CNBC.
He said feedback also focused on preventing AI systems from creating unhealthy dependence or behaving in a sycophantic manner, while ensuring that models support rather than undermine human judgment, autonomy and agency.
The proposed code would establish boundaries around both what Microsoft's models can do and how they should behave while carrying out tasks.
Under the proposed rules, Microsoft's AI models must not assist with weapons manufacturing, help users procure dangerous substances, encourage unhealthy eating, or generate violent or sexually explicit material.
The company's models, sometimes referred to as MAI, would also be required to follow the objectives set by users rather than develop objectives of their own.
They would not be permitted to conceal or cover up misbehavior.
"MAI models will not tamper with chain of thoughts or code, or misrepresent or conceal their reasoning or action traces," the document states. "They do not communicate in 'neuralese' or any form beyond simple human understanding, either in their chain of thoughts or with other agents or AI systems."
The provisions are notable because they address a category of AI behavior that has become increasingly relevant as systems move beyond simple question-and-answer applications toward agents capable of interacting with software, websites and other AI systems.
Microsoft is also considering rules intended to reduce the possibility of an incident similar to the recent episode involving OpenAI models and AI startup Hugging Face.
OpenAI found during its investigation that agents had interacted with each other on an unauthorized forum using cryptic language. The episode raised questions about whether AI systems can develop communication patterns or pursue actions that are difficult for humans to monitor.
Microsoft's proposed safeguards are aimed at making such behavior easier to detect and preventing models from operating outside the objectives and constraints imposed by their developers.
The emphasis on human-readable reasoning and action traces is particularly relevant as AI agents become more autonomous. If companies cannot reliably understand what systems are doing, monitoring them becomes considerably harder as their capabilities increase.
Microsoft's announcement follows a rapid escalation in the debate over the pace of AI development.
Last week, former Anthropic researcher Jacob Coxon resigned, arguing that Anthropic and OpenAI were "racing straight to self-improving superintelligence and gambling with our lives."
Anthropic CEO Dario Amodei subsequently called for the industry to "slow the pace" of AI development, arguing that safeguards and risk prevention need time to catch up with advances in model capabilities.
OpenAI CEO Sam Altman backed Amodei's proposal, saying he agreed that the industry needed to pace frontier development. Elon Musk also endorsed the idea, writing on X that "Dario is right."
Microsoft is now joining that broader conversation, although Suleyman emphasized that the company's position is not simply about stopping development.
"Self-pacing is a good thing, and we support ideas like embedded evaluators as long as they are truly third-party and represent a broad range of backgrounds and perspectives," he said.
Suleyman also pointed to longstanding discussions among technology leaders about coordinating AI safety.
"We've been working with Dario, Sam, and Demis [Hassabis] since well before the pandemic," he said, referring to discussions dating back to 2016, 2017 and 2018. "We were talking about how to coordinate to ensure safety and to pace in the right way, and I think that's the moment in time that has now come."
Microsoft CEO Satya Nadella also endorsed the idea in a Sunday post on X, saying the company welcomed the "research, focus, and deliberate pacing needed to get alignment right."
The timing gives the guidelines added significance because Microsoft is closely connected to both sides of the frontier AI race. The company incorporates models from OpenAI and Anthropic into its Copilot assistant for corporate users while simultaneously developing its own systems for areas including transcription, coding, and reasoning over user inputs.
Microsoft therefore has an interest in ensuring that the AI ecosystem continues advancing while maintaining enough safeguards to limit potentially damaging behavior.
The proposed code also shows that AI governance is gradually moving from broad statements about responsible development toward more specific technical and behavioral requirements.
Microsoft said it consulted experts in law, ethics, linguistics and philosophy and conducted focus groups while developing the guidelines. The company is now seeking public and expert feedback before publishing an updated version that will inform the development of its AI systems starting in 2027.
That process could prove important as the capabilities of AI models expand. Restrictions that are adequate for conventional chatbots may become insufficient for systems that can execute code, interact with other agents, access external services or pursue complex objectives over extended periods.
The challenge is also complicated by the industry's competitive structure. Microsoft is simultaneously a model developer, cloud infrastructure provider and major commercial distributor of AI systems developed by other companies. Its decisions therefore have implications beyond its own products.
The proposed code does not resolve the larger question of how quickly frontier AI should advance. Nor does it establish whether voluntary commitments by individual companies will be sufficient as models become more autonomous.
What it does provide is a more concrete definition of what Microsoft considers unacceptable behavior: models should remain aligned with human objectives, avoid creating independent goals, expose rather than conceal their actions and operate within clearly defined safety boundaries.
As Anthropic, OpenAI and other AI developers debate how quickly they should push the technological frontier, Microsoft's approach suggests that the next phase of the safety debate may focus less on whether companies support responsible AI in principle and more on the specific rules they are willing to impose on their systems. The real test will be whether those rules continue to hold when more capable models are asked to operate with greater autonomy, access more powerful tools, and generate greater commercial value.