Gemini 4 Argon Leads 13 of 18 Benchmarks Google Chose, but Only Cyber Defenders Can Use It
A.I. / news
Gemini 4 Argon Leads 13 of 18 Benchmarks Google Chose, but Only Cyber Defenders Can Use It
Google priced the model at $2 and $10 per million tokens and gave access first to its Fairwind Program, with a guardrail-free version for trusted defenders.

Google announced Gemini 4 Argon on Sept. 30, a model it says leads or ties on 13 of the 18 benchmarks it disclosed, and it is available only to vetted cyber defenders for now.
Koray Kavukcuoglu, senior vice president of Google DeepMind and Chief AI Architect at Google, wrote the announcement on Google's blog. Access begins with Google's Fairwind Program, which the post describes as free cybersecurity defence for critical public infrastructure, run with security firms such as Wiz.
Google said broader availability will follow "as soon as possible", starting with paid API customers and Google AI Ultra subscribers. It gave no date.
Gemini 4 Argon pricing and output limit
Introductory pricing is $2 per million input tokens and $10 per million output tokens. Google said the standard rate afterwards is $4 and $20. Cached input tokens carry a 95% discount off the input price.
The output limit is 1 million tokens, up from 64,000 on the previous model, according to the post. Google did not publish a context window figure in the material reviewed.
One day earlier OpenAI priced GPT-6.1 Sol at the same $2 and $10. OpenAI is also the subject of an FTC consumer-risk investigation alongside Anthropic, the two rivals Google benchmarked against.
The benchmarks are Google's own
VentureBeat's Carl Franzen tabulated the comparison against OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5. The figures below are vendor-supplied. VentureBeat lists Google as the source for each, and no independent party is named as having run them.
| Benchmark | Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| AutomationBench | 51.3% | 41.4% | 42.5% |
| Vals Finance Agent v2 | 65.4% | 53.5% | 58.6% |
| LVBench | 91.7% | 87.5% | 83.7% |
| CWE-bench | 68% | 68% | 67% |
The table is a selection. Argon leads outright on 12 of the 18 disclosed benchmarks and ties on one. GPT-6 Astra leads on three, including FrontierSWE v2 at 65.5% against Argon's 55.0%. Claude Opus 5.5 leads on two, including Terminal-bench 4.0 at 66.4% against 57.4%.

- Gemini 4 Argon77.9 %
- Claude Opus 5.574.2 %
- GPT-6 Astra74.1 %
Source: VentureBeat, Sept. 30, 2026, citing Google's benchmark table
The gap on the software-engineering test is 3.7 points over the nearest rival. On CWE-bench, the vulnerability-finding test closest to the model's launch audience, Argon ties Astra at 68%.
Why the guardrails are off for defenders
Google said Argon is "designed to refuse harmful requests while preserving legitimate, dual-use scientific research". For Fairwind participants it will release a version without cyber guardrails.
The post says Argon "can autonomously find, validate, and patch critical software vulnerabilities". That is Google's claim. The post does not say how many defenders have access, how they are vetted, or what stops the guardrail-free build from leaving the programme.
Google did not say whether the standard model will ship with the same refusal behaviour as the Fairwind version.
The company said wider rollout will begin with paid API customers and Google AI Ultra subscribers. Until then, no customer outside Fairwind can verify any of the 18 figures. The same gap applies to OpenAI's chip-design model deal with Synopsys, which named no benchmarks at all.
Sources
More in A.I.
- 01Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.
- 02OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.
- 03Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2BAWS's Strands Labs scores 72.3 percent on JevBench at a 106 ms median on an RTX 3090. TypeSafe's CEO calls the current crop of rivals less serious than his own team.
- 04OpenAI and Synopsys Sign Chip-Design Model Deal With No Customers or Benchmarks NamedGPT-Synopsys will run Synopsys EDA tools on OpenAI-hosted infrastructure under a revenue-sharing agreement. The Sept. 30 announcement gives no dollar figure and no ship date.