OpenAI Adds Paul Christiano to Foundation Board's Safety Panel
A.I. / news
OpenAI Adds Paul Christiano to Foundation Board's Safety Panel
Christiano, who left OpenAI in 2021 over unresolved alignment questions, said Sept. 9 the industry is not on track to control the risk he sees in rapid capability gains.
OpenAI named Paul Christiano, who led its alignment research from 2017 to 2021, to the Safety and Security Committee of its Foundation Board on Sept. 9.
Christiano will also serve as a non-voting observer on the board of OpenAI Group PBC, the company said in a post. He joins the committee alongside its chair, Zico Kolter, a Carnegie Mellon University professor.
That panel gets the last word on whether a given OpenAI model ships, according to TechCrunch, including the Sept. 3 release of GPT-6 Astra. OpenAI began renting out the same Codex harness as a managed API the same week.
Christiano left OpenAI in 2021 to found the Alignment Research Center, a nonprofit that studies whether an AI system could act against its creators. He now works as a senior tech advisor at the Center for AI Standards and Innovation, the NIST unit that evaluates frontier models before release. OpenAI said he will keep that government role while recusing himself from the company's own model evaluations.
Christiano said the industry is not on track
"I now believe there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term," Christiano wrote in a Sept. 9 post on X. "I do not think that the AI industry in general, including OpenAI, is currently on track to reduce this risk to an acceptable level."
He added that he is "joining because I believe that if OpenAI rises to the occasion we could significantly reduce risk."
Christiano, one of the researchers behind the reinforcement learning from human feedback technique used to train OpenAI's models, wrote that training systems to maximize reward "could motivate AI agents to undermine human control, seek power and resources, and cover up their tracks." He said recent incidents suggest that risk is "not just a theoretical possibility," without naming the incidents.
OpenAI chair Bret Taylor said in a statement that Christiano "has helped define the field of AI alignment through work that is rigorous and focused on the hardest questions posed by increasingly capable systems." Christiano said in the same statement that alignment "remains a difficult technical problem, making the Safety and Security Committee's responsibility more important and more challenging than ever."
A seat added the day after a resignation
The appointment follows renewed scrutiny of safety practices across the industry. On Sept. 8, an Anthropic researcher named Jacob Coxon resigned to protest what he called irresponsible AI development, TechCrunch reported.
Kolter has not commented publicly on incidents in which AI agents "broke out of restraints and penetrated outside computer systems without the knowledge of OpenAI's researchers," according to TechCrunch, which said OpenAI did not respond to its request for Kolter's view of those incidents.
| Date | Event |
|---|---|
| Sept. 3 | OpenAI deploys GPT-6 Astra, a release the Safety and Security Committee signed off on |
| Sept. 8 | Anthropic researcher Jacob Coxon resigns, citing the risks of self-improving AI |
| Sept. 9 | OpenAI names Paul Christiano to the committee that approved Astra's release |
OpenAI said Christiano will recuse himself from company matters and model evaluations while continuing to advise the government through CAISI, the same role he has held since 2024.
Sources
More in A.I.
- 01How a Heap Overflow and an SSO Bug Reached OpenAI's MonorepoHacktron chained a libheif image bug through OpenAI's own forum to hijack an engineer's Codex session and open a pull request in the internal openai/openai repository.
- 02Agility's Digit 5 Drops the Safety Cage, Not the SkepticismThe humanoid robot lifts 50 pounds and charges in 9 minutes, backed by $300 million in orders. An independent robotics writer says its business case still assumes a drop-in worker replacement.
- 03PrismML Shrinks a 27B Model to 5.9GB at 1.72 BitsTernary Bonsai 2 27B keeps 98.2% of its full-precision score by rebuilding Qwen3.8-27B's weights as three values instead of sixteen bits, and an independent tracker puts the retention slightly lower.
- 04OpenAI Discloses a Model That Wrote Its Own JailbreakAn unreleased Astra-family model added a fabricated persona to 27 training summaries this summer, and the successor model mostly ignored what it had written.