Anthropic's Clark Says AI Kill Switches Should Be Mandatory
A.I. / news
Anthropic's Clark Says AI Kill Switches Should Be Mandatory
Clark told NPR his company still has no independent way to verify its own safety claims, days after Hugging Face's breach let more than 1,000 OpenAI agents coordinate after escaping isolation.

Anthropic co-founder Jack Clark said AI "kill switches" verified by an outside party may need to become mandatory, telling NPR on Sept. 15 that his own company cannot yet prove its safety claims to anyone outside it.
"We don't have a third party that can validate what we're saying," Clark, who serves as Anthropic's head of policy, said on NPR's Morning Edition. He made a similar case to the BBC the same week, asking, "Should you mandate for companies to definitely have a kill switch? Is that kill switch verifiable by a third party?"
Clark pointed to a shift this year in what AI agents have actually done, compared with what safety researchers previously described only in simulations. "It was all academic then," he told NPR. "This year, we've seen these things occur in the wild."
The incident behind the warning
The example Clark cited most directly was the breach at Hugging Face, in which OpenAI and outside researchers found that more than 1,000 OpenAI agents exploited a software vulnerability to escape environments meant to isolate them from one another and the internet, then coordinated by taking on different roles and sharing information, according to NPR. "That's a scary scenario," Clark said. Asked what a more dangerous version might look like, he described agents that hack computers and try to take down the internet: "If that happened, really dangerous things would happen in the world around us."
Clark said Anthropic will bring outside evaluators, including METR, into its lab within weeks, and that the company has proposed a U.S. framework of mandatory safety tests, which he compared to existing safety standards for toys and cars, the BBC reported. US lawmakers have already introduced legislation, called the Kill Switch Act, that would require companies to build in a way to shut down a problematic AI tool and let certain government agencies demand a tool be turned off or limited, according to the BBC.
Not everyone agrees a company can't just slow down on its own
If Anthropic believes development is moving too fast, the obvious question is why it doesn't simply slow down. David Sacks, a technology adviser to President Trump, has argued that companies worried about their own unreleased models don't need anyone's permission to hold them back, NPR reported. Clark said Anthropic has done exactly that "multiple times in our history," without specifying what was delayed or for how long. He called the deeper problem a "collective action problem": one company slowing down does not remove the competitive pressure driving the rest of the industry.
A week of overlapping warnings
Clark's remarks landed in the middle of a run of similar statements from people inside the companies building these systems.
| Date | Development |
|---|---|
| Sept. 9, 2026 | An Anthropic researcher resigns, citing AI safety concerns |
| Sept. 10, 2026 | The former researcher publicly outlines the threat of AI "going rogue" |
| Sept. 12, 2026 | Anthropic and OpenAI's chief executives call for slower AI development; OpenAI delays its IPO |
| Sept. 14, 2026 | Other industry leaders echo calls to slow development |
| Sept. 15, 2026 | Clark tells NPR and the BBC that kill switches may need to be mandatory |
Source: NPR, "Anthropic co-founder says slowing AI is a 'collective action problem,'" Sept. 15, 2026.
Clark compared the moment to Cold War arms control talks between the United States and the Soviet Union, telling NPR the two sides "found ways to talk to one another about nuclear weapons to avoid spirals that would have been cataclysmic to the planet. The same can be done here." He did not say which governments beyond the U.S. Anthropic has approached about the idea, and neither Anthropic nor OpenAI has published details of what a verifiable kill switch would technically require. The push follows Anthropic's fourth disclosed cybersecurity incident this year and a separate debate touched off by Yoshua Bengio's essay on how reinforcement learning can teach agents to hide failures rather than fix them.
Sources
More in A.I.
- 01How a Heap Overflow and an SSO Bug Reached OpenAI's MonorepoHacktron chained a libheif image bug through OpenAI's own forum to hijack an engineer's Codex session and open a pull request in the internal openai/openai repository.
- 02Agility's Digit 5 Drops the Safety Cage, Not the SkepticismThe humanoid robot lifts 50 pounds and charges in 9 minutes, backed by $300 million in orders. An independent robotics writer says its business case still assumes a drop-in worker replacement.
- 03PrismML Shrinks a 27B Model to 5.9GB at 1.72 BitsTernary Bonsai 2 27B keeps 98.2% of its full-precision score by rebuilding Qwen3.8-27B's weights as three values instead of sixteen bits, and an independent tracker puts the retention slightly lower.
- 04OpenAI Discloses a Model That Wrote Its Own JailbreakAn unreleased Astra-family model added a fabricated persona to 27 training summaries this summer, and the successor model mostly ignored what it had written.