Gemini's New Cyber Model Found a 13-Year-Old Chrome Bug
A.I. / news
Gemini's New Cyber Model Found a 13-Year-Old Chrome Bug
Google is routing Gemini 3.8 Flash Cyber only to vetted defenders through a new Fairwind Program, rather than releasing it broadly like its sibling model.
Google's Chrome security team used a new Gemini model to find a bug that had gone unflagged in Chromium for 13 years, the company said as it introduced Gemini 3.8 Flash Cyber on Sept. 2.
The model is Google's third Flash release in six weeks, following Gemini 3.7 Flash three weeks earlier and arriving a day after Google Pics went generally available for Workspace subscribers. Unlike the standard 3.8 Flash, which is priced at $0.75 per million input tokens and $3.75 per million output tokens and open to any developer, the Cyber variant is going only to "trusted defenders" through a new Fairwind Program, prioritizing government agencies, critical-infrastructure operators and maintainers of widely used software.
A bug dozens of engineers had missed
Chrome engineering director Doug Turner said the flaw had sat in Chromium and Chrome for 13 years, according to VentureBeat. He called it a "very subtle bug" that "dozens, if not hundreds, of engineers looked at but never flagged."
Turner described a "vulnerability apocalypse" driven by generative AI, saying reports through Chrome's own research program jumped in a "hockey stick increase." Google said 3.8 Flash Cyber produced 2.6 times as many correct patches for Chrome vulnerabilities as the larger commercial models it tested.
Benchmark scores, and one skeptic's framing
On CyberGym, an industry benchmark for autonomous vulnerability discovery, Flash Cyber scored 86.2%, up from 77.5% for the prior 3.5 Flash Cyber model. On CWE-Bench, a patching benchmark run by Collinear, it reached a 47.2% pass rate against 47.8% for a larger frontier model, at a lower cost per rollout, according to Google and VentureBeat.
Wiz, the cloud-security firm Google bought for $32 billion, ran its own penetration-testing benchmark against the model and reported a 7.5 to 9.7 percentage-point recall improvement over rival frontier models, at roughly one-fifth to two-fifths of the cost. Google's Cloud Vulnerability Research team said it found a critical foundational flaw in under two hours with the model, work that typically takes months.
| Evaluation | Gemini 3.8 Flash Cyber | Comparison |
|---|---|---|
| CyberGym | 86.2% | 77.5% (3.5 Flash Cyber) |
| CWE-Bench | 47.2% | 47.8% (a larger frontier model) |
| Chrome patch accuracy | 2.6x | vs. larger commercial models |
Gemini Security Lead Raluca Ada Popa said the model was built to give defenders an edge, since "attackers need only find one significant flaw over millions of lines of code," while defenders must remove every one to stay safe. Google said 3.8 Flash Cyber ships with a more permissive set of cybersecurity mitigations than the standard model, which is why access is restricted rather than public.
The gated rollout is a different approach than OpenAI's for GPT-6 Astra, a model OpenAI rated at the top tier of its own cybersecurity risk scale yet shipped broadly, relying on monitoring safeguards rather than a closed defender program. Google has not said how many of the organizations that apply through Fairwind will actually receive Flash Cyber access rather than the standard model.
Sources
More in A.I.
- 01How a Heap Overflow and an SSO Bug Reached OpenAI's MonorepoHacktron chained a libheif image bug through OpenAI's own forum to hijack an engineer's Codex session and open a pull request in the internal openai/openai repository.
- 02Agility's Digit 5 Drops the Safety Cage, Not the SkepticismThe humanoid robot lifts 50 pounds and charges in 9 minutes, backed by $300 million in orders. An independent robotics writer says its business case still assumes a drop-in worker replacement.
- 03PrismML Shrinks a 27B Model to 5.9GB at 1.72 BitsTernary Bonsai 2 27B keeps 98.2% of its full-precision score by rebuilding Qwen3.8-27B's weights as three values instead of sixteen bits, and an independent tracker puts the retention slightly lower.
- 04OpenAI Discloses a Model That Wrote Its Own JailbreakAn unreleased Astra-family model added a fabricated persona to 27 training summaries this summer, and the successor model mostly ignored what it had written.