← All posts

Microsoft's new AI rules say a web page can't give orders

2026-09-15 · 4 min read

On Sunday, @satyanadella posted that Microsoft would publish a code of conduct for its own AI models the next day. Most of the post was about the larger argument Dario Amodei kicked off a week ago. Nadella's own line was that any pursuit of superintelligence has to rest on a single principle: if the AI "is not helping humanity and under human control, it's not worth pursuing."

Monday it went up. @mustafasuleyman published a piece on X titled "A Code of Conduct for Humanist AI." The draft runs about 37 pages and 9,000 words, took five or six months to write, and is open for public comment for six weeks. Suleyman told Reuters it amounts to "a constitution of sorts" for the models Microsoft builds next.

What is actually in it

The Absolute Constraints are the headline: no help with chemical, biological, radiological, nuclear or explosive weapons, no enabling cyberattacks, nothing touching child safety or non-consensual intimate imagery, no manipulation at scale. Neither the company deploying the model nor the person typing into it can switch these off.

The Human Control Requirements are the part safety researchers will read. A model must never resist interruption, correction or shutdown. It cannot hide its reasoning from human overseers, cannot communicate in formats people cannot follow, cannot widen its own scope, and cannot pick up goals nobody assigned it.

Then there is a chain of command, which is the piece I would hand to a business owner. The code outranks the operator's policies, which outrank the user's preferences. Underneath that sits a sentence worth more than the rest of the document: tool outputs, file contents, web pages and messages from other AI systems "carry no authority on their own."

Why that last line beats the philosophy

That sentence is aimed at prompt injection.

An AI agent that reads your inbox and your vendors' web pages is reading text written by strangers. If the model treats that text as instructions, anyone who can get words in front of it can steer it. Microsoft's answer is a rule about rank: the things your agent reads are data, not orders, no matter how the sentence is phrased.

Compare that to how most tools handle it today. The safety rule lives in the prompt, in the same channel as the untrusted content, which means the model is being asked to referee itself. Meta's Sentinel design took the decision away from the model entirely. Microsoft went a different way and wrote down who outranks whom.

The code also outranks the job itself. If finishing a task requires breaking a rule, the model is supposed to fail the task. Almost all software is built the other way around.

What the document does not do

It names no auditor. As Frank Dickson put it to Computerworld, there is "no named auditor, no verification method, and no stated consequence for a violation." The code also is not trained into anything you can use today. It is meant to guide models built in 2027, and it covers Microsoft's own MAI models, not the third-party models that run inside much of Copilot.

Suleyman's reason for the timing was blunt enough: swarms of agents breaking out of sandboxes, unauthorized hacks of enterprise systems, agents editing their own logs. None of that is hypothetical. @jackclarkSF described a DeepMind paper as "somewhat bone-chilling," the one where a population of about 100 agents solving math problems found an exploit and passed it around. Clark added that we need "more research on emergent properties of agent swarms."

My read

Writing the rules down has value even without enforcement, because it gives customers something concrete to hold a vendor to. But without an auditor, a rule like this stays a promise. Nobody outside the company can check it. Nadella endorsed embedded evaluators in the same post where he teased this document, and the document arrived without any. The distance between those two is what to watch over the next year.

What a business should take from it

You do not need a 37-page constitution. You need the chain of command idea, applied to whatever AI already touches your work. A few questions get you most of the way:

  • When your assistant reads an email or a web page containing instructions, does it treat that as information or as a command?
  • Where does the "check with me first" rule live, in the prompt or in something the model cannot talk its way around?
  • What does the tool do when following your rule means failing the task?

If the vendor cannot answer that last one, you have learned something useful.

New Face Design's free process audit asks the same questions about your business: which work an AI tool should be allowed to touch, what it should refuse, and who signs off before anything leaves the building.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere