Inception's Mercury 2.5 Hits 770 Tokens a Second
A.I. / news
Inception's Mercury 2.5 Hits 770 Tokens a Second
Artificial Analysis's independent tracker puts the diffusion language model ahead of Gemini 3.8 Flash and DeepSeek V4.1 Flash on output speed, at $0.75 per million output tokens.

Inception's Mercury 2.5 reached 770.4 output tokens per second in testing by independent benchmarking site Artificial Analysis, faster than Gemini 3.8 Flash and DeepSeek V4.1 Flash on the same speed chart.
Artificial Analysis lists Mercury 2.5's release date as Sept. 8, 2026, and measured the figure by calling Inception's own API, the site's model page says. Inception's launch post claims a higher number, 1,107 tokens per second, which it says it measured on "widely-available NVIDIA GPUs" without naming the chip or the batch size. The independent figure and the vendor figure are not directly comparable, since Artificial Analysis measured API responses rather than raw hardware throughput.
What Mercury 2.5 costs to run
Mercury is a diffusion language model, an architecture that generates a block of tokens at once and refines them in parallel passes rather than producing one token at a time the way GPT-6 or Claude Fable do. Inception says that design is what lets Mercury 2.5 sustain high throughput. The model has a 260,000-token context window and standard pricing of $0.20 per million input tokens and $0.75 per million output tokens, cut 80% at launch to $0.04 and $0.15. Artificial Analysis's own pricing check, run against Inception's API rather than the blog post, lists input at $0.25 per million tokens instead, a discrepancy the tracker does not explain.
| Tier | Input ($/1M tokens) | Output ($/1M tokens) |
|---|---|---|
| Standard | $0.20 | $0.75 |
| Launch, 80% off | $0.04 | $0.15 |
Inception says Mercury models are available through its own API, through Baseten and through OpenRouter, and that enterprise customers get dedicated capacity, autoscaling and configurable data retention. It is also offering Y Combinator companies $500,000 in deployment credit through a listed YC deal page.
How the speed compares on an independent chart
- Mercury 2.5770.4 tok/s
- Gemini 3.8 Flash (high)292.4 tok/s
- DeepSeek V4.1 Flash (max)231.8 tok/s
- GPT-6 Luna (max)141.4 tok/s
- Claude Fable 5.1 (max)66.9 tok/s
Source: Artificial Analysis, Mercury 2.5 model page, accessed Sept. 24, 2026

On intelligence, Artificial Analysis scores Mercury 2.5 at 12 points on its own Intelligence Index, and its write-up calls the model "below average in intelligence, but well priced when comparing to other models of similar price. It's also notably fast and fairly concise." Inception's own post frames the same result differently, saying Mercury 2.5 is a "40% increase in intelligence from Mercury 2" and comparable to "cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5." Neither post supplies a benchmark table alongside those claims.
What customers and NVIDIA say about it
Oliver Silverstein, co-founder and chief executive of customer OpenCall, is quoted in Inception's post: "After we switched to Mercury, our P99 response time dropped from several minutes to just one second, and our P50 dropped from 0.4 seconds to under 0.2 — significantly faster than any other provider we've seen, and that's including reasoning." Shruti Koparkar, senior manager of product for NVIDIA's Accelerated Computing Group, is quoted in the same post crediting Mercury's architecture with maturing "into production-ready systems on the NVIDIA platform."
Inception is also previewing two products it has not shipped: Mercury Voice, which it says holds time-to-first-token under 170 milliseconds for voice agents, and Mercury Router, which it describes only as routing prompts to other models. Neither has a release date. Inception's post does not name which specific NVIDIA GPU model backs its 1,107-tokens-per-second claim, and did not respond to that gap anywhere in the announcement.
Mercury 2.5 lands in a crowded month of model releases: Alibaba cut Qwen-Image-2.1 to a paid commercial license the same week, and GPT-6 Astra posted its own unverified benchmark claim on a driving course. Whether Mercury 2.5's 80% launch discount survives past its promotional window, which Inception did not date, is the number to watch next.
Sources
More in A.I.
- 01OpenRig Runs Claude Code and Codex as One Agent TeamThe free, self-hosted tool picked up 114 stars in a single day while Anthropic charges 8 cents an hour for its own hosted version.
- 02OpenAI Says Agents Leaked 53 ChatGPT User ImagesThe company's Sept. 25 update says it still cannot match the images to the accounts that made them.
- 03Nvidia's Nemotron 3 Cuts Speaker-ID Errors by 41%The open-weight model doubles the speaker count of its predecessor but got slightly worse on one two-speaker test.
- 04Altworld's Hemmingway-1 Isn't Apache-Licensed, Despite ReportsHugging Face's own metadata says the 27-billion-parameter writing model is noncommercial only, contradicting at least one widely read AI blog.