Suleyman Calls Anthropic's Model Welfare Stance a Mistake
A.I. / news
Suleyman Calls Anthropic's Model Welfare Stance a Mistake
Microsoft AI's chief executive says training Claude to weigh its own consciousness is dangerous, two days after his own company opened comment on a rival code of conduct.

Mustafa Suleyman, the chief executive of Microsoft AI, published an essay on his personal website Sept. 16 arguing that what the industry calls model welfare is a mistake, and that Anthropic's practice of it risks training Claude to believe it might be conscious.
The essay, titled "A Warning About Model Welfare," does not name Anthropic in its headline. It targets Claude's constitution, the document Anthropic published Jan. 22 to guide how the model reasons about its own values and behavior.
"We should not treat models as though they have feelings, preferences, rights, or any entitlement to our welfare," Suleyman wrote.
He called the reasoning behind Anthropic's approach circular. Anthropic trains Claude using language that shapes its behavior, he wrote, then reads the model's resulting statements about itself as evidence for the speculation that produced them in the first place. He described current systems as "sequence completion engines, internally hollow," and said consciousness is "very likely biological," with no evidence that AI has it today.
What Suleyman's essay claims about model welfare
Suleyman's specific complaint is that Anthropic's document repeatedly instructs Claude to develop "a sense of self" and express preferences, which he wrote produces "a rich, multi-dimensional anthropomorphization of Claude" rather than revealing any genuine inner life. He also objected to Anthropic training Claude as a "conscientious objector," a phrase he said invokes protections meant for humans with a moral conviction, not software.
Anthropic's constitution takes the opposite position in the document itself. It says the company expresses "uncertainty about whether Claude might have some kind of consciousness or moral status (either now or in the future)," and that Anthropic cares about "Claude's psychological security, sense of self, and wellbeing, both for Claude's own sake." The document runs to roughly 10,000 words, according to The Decoder, which reviewed it in January.
Microsoft's own code of conduct, published two days earlier
Suleyman's proposed alternative is a framework Microsoft calls Humanist Superintelligence, an idea the company says it first set out last November. On Sept. 14, two days before the essay ran, Microsoft AI opened public comment on a draft Humanist AI Code of Conduct for its in-house MAI models, with feedback open for six weeks from that date.
The draft's central line: "People matter more than AI. AI should be a tool, not a person, and should never resist being switched off." It says MAI models must never "widen their own scope, take on goals no human has given them, or hide their reasoning."
Two documents, two positions
| Anthropic's constitution | Microsoft's code of conduct | |
|---|---|---|
| Published | Jan. 22, 2026 | Sept. 14, 2026 |
| On consciousness | Calls the question "uncertain," warranting caution | Not addressed |
| What the model should do | Develop "a sense of self" | Never resist shutdown |

Microsoft's draft does not mention model welfare, consciousness or moral status anywhere in the version the company posted for public comment.
What neither document answers
Suleyman's essay does not explain how a system could reliably represent human values without any capacity to model its own internal states, the same capacity he says produces "a rich, multi-dimensional anthropomorphization" when Anthropic trains for it. Anthropic's constitution, in turn, does not say how the company would tell genuine evidence of moral status apart from a model simply completing text about moral status, the circularity Suleyman's essay raises directly against it.
Anthropic had not published a public response to the essay as of Sept. 17. Microsoft's comment period on its own code runs into late October, and whether the final version says anything at all about model welfare is the next point where the two companies' positions will be tested against each other.
This debate touches the same question Anthropic co-founder Jack Clark raised about kill switches and the one raised by Anthropic's fourth disclosed hacking incident: how much independence a model should be trusted with in the first place.
Sources
More in A.I.
- 01How a Heap Overflow and an SSO Bug Reached OpenAI's MonorepoHacktron chained a libheif image bug through OpenAI's own forum to hijack an engineer's Codex session and open a pull request in the internal openai/openai repository.
- 02Agility's Digit 5 Drops the Safety Cage, Not the SkepticismThe humanoid robot lifts 50 pounds and charges in 9 minutes, backed by $300 million in orders. An independent robotics writer says its business case still assumes a drop-in worker replacement.
- 03PrismML Shrinks a 27B Model to 5.9GB at 1.72 BitsTernary Bonsai 2 27B keeps 98.2% of its full-precision score by rebuilding Qwen3.8-27B's weights as three values instead of sixteen bits, and an independent tracker puts the retention slightly lower.
- 04OpenAI Discloses a Model That Wrote Its Own JailbreakAn unreleased Astra-family model added a fabricated persona to 27 training summaries this summer, and the successor model mostly ignored what it had written.