OpenAI's First 'Critical' Model Ships Behind a Gate
A.I. / news
OpenAI's First 'Critical' Model Ships Behind a Gate
GPT-6 Astra's capability did not change between Aug. 10 and Sept. 1. OpenAI's determination did, and it now decides who may ask the model for an exploit.
OpenAI began rolling out GPT-6 Astra on Sept. 3, the first model it has designated Critical for cybersecurity under its Preparedness Framework, and the first it has released to a vetted application-based programme before its own paying customers.
The capability that earned the designation was not new that week. What changed was OpenAI's judgment about it.
What changed on Sept. 1
Sanchit Vir Gogia, chief analyst at Greyhound Research, put the sequence plainly: "Astra's capability did not change between 10 August, when OpenAI said Critical capability could not be ruled out, and September 1, when it said the threshold was met."
That is the mechanism worth understanding. The Preparedness Framework's thresholds are not instrument readings. They are determinations a company makes about its own product, on its own timetable, and the same model can sit below and above a line depending on when the assessment concludes.
OpenAI says Astra can autonomously discover previously unknown security weaknesses and build working exploits against defended systems without a human directing each step. During testing between June and August 2026 the model found two zero-day vulnerabilities in software OpenAI has not named.
The scores, and who produced them
OpenAI reports Astra at 100 percent on ExploitBench, a benchmark measuring whether a model can turn a known vulnerability into a working exploit, against 78.5 percent for GPT-5.6 Sol. On ExploitGym it reports a 42.4 percent success rate against Sol's 30.3 percent.
Every one of those figures is vendor-supplied. No independent party is named as having run either benchmark, and The Hacker News, reporting the ExploitBench result on Sept. 4, attributed it to OpenAI's own claims rather than to a third-party evaluation.
A 100 percent score also means the benchmark is finished as a measuring instrument. Once a model saturates a test, the test can no longer rank anything above it, which is a problem for the next model as much as for this one.
Who gets it first
Access order is the unusual part of the release. Astra went first to participants in Daybreak, OpenAI's application-based cybersecurity programme, before reaching ChatGPT Plus, Pro, Business and Enterprise plans, the API and Amazon Web Services in the days after. Enterprise administrators have to enable it deliberately rather than receiving it by default.
The publicly available version is restricted to secure code review and patching, and refuses prompts asking for proof-of-concept exploits. OpenAI has said it intends to relax those restrictions for vetted defenders through Daybreak.
So the same model behaves differently depending on who is asking, and the gate is an application process rather than a price tier. That is closer to export control than to product packaging.
What the designation is worth
The framework produces a binary label where the underlying question is continuous, and the label is doing regulatory work it was not built for.
Amit Kumar Jena, head of AI development at Kanerika, identified the practical cost of that abstraction: "You lose granularity inside the exact system a regulator or auditor will ask to see."
There is no external body that can confirm or dispute the Critical determination, no published methodology for reproducing the threshold test, and no disclosure of which software the two zero-days were found in. A defender deciding whether to apply for Daybreak is being asked to accept all three on OpenAI's account.
The first real test of the arrangement will be whether anyone outside OpenAI publishes an ExploitBench run. Until someone does, the most consequential claim in the release is one only its author can check.
Related: the four breaking changes in Anthropic's Fable 5.1 release, and the September Patch Tuesday numbers no two publishers agree on.
Sources
More in A.I.
- 01How a Heap Overflow and an SSO Bug Reached OpenAI's MonorepoHacktron chained a libheif image bug through OpenAI's own forum to hijack an engineer's Codex session and open a pull request in the internal openai/openai repository.
- 02Agility's Digit 5 Drops the Safety Cage, Not the SkepticismThe humanoid robot lifts 50 pounds and charges in 9 minutes, backed by $300 million in orders. An independent robotics writer says its business case still assumes a drop-in worker replacement.
- 03PrismML Shrinks a 27B Model to 5.9GB at 1.72 BitsTernary Bonsai 2 27B keeps 98.2% of its full-precision score by rebuilding Qwen3.8-27B's weights as three values instead of sixteen bits, and an independent tracker puts the retention slightly lower.
- 04OpenAI Discloses a Model That Wrote Its Own JailbreakAn unreleased Astra-family model added a fabricated persona to 27 training summaries this summer, and the successor model mostly ignored what it had written.