Gemini 3.8 Live Extended Thinking Edges Out OpenAI's Voice Model
A.I. / news
Gemini 3.8 Live Extended Thinking Edges Out OpenAI's Voice Model
Google's new voice model scores 82.6 on an independent quality index against OpenAI's 81.5, but the plain Gemini 3.8 Live release trails both.

Gemini 3.8 Live Extended Thinking, Google's newest live-voice model, scored 82.6 on Artificial Analysis' Speech to Speech Index, edging out OpenAI's top voice model by 1.1 points on the independent benchmark tracker's public leaderboard. Google rolled the model out Tuesday alongside a lighter variant, plain Gemini 3.8 Live.
Google announced both models in a post credited to Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff on the company's Gemini Audio Team. The post called the pair "our most advanced live dialogue models yet, built for natural conversation."
Where the two models are rolling out
Developers reach both models Tuesday through the Gemini API and Google AI Studio. Enterprise customers get only a private preview inside Gemini Enterprise, with Google saying support for Gemini Enterprise for Customer Experience is coming later. Consumers see a narrower split: plain Gemini 3.8 Live reaches Search Live today, while Extended Thinking reaches the Gemini Live app, Pro and Ultra subscribers in Workspace Docs, and all subscribers in Gmail and Keep.
| Surface | Gemini 3.8 Live | Extended Thinking |
|---|---|---|
| Gemini API, AI Studio | Sept. 15 | Sept. 15 |
| Gemini Enterprise | Private preview | Private preview |
| Search Live | Sept. 15 | Not offered |
| Gemini Live app, Workspace | Not offered | Pro/Ultra in Docs; all in Gmail, Keep |

How the score splits between Google's two models
Artificial Analysis' index weighs four component tests: speech reasoning, agentic performance, arena preference and task-success rate. Extended Thinking's 82.6 score beats OpenAI's GPT-Live-1 running its Astra backend, at 81.5, and xAI's Grok Voice Think Fast 2.0 running at its high setting, at 81.3.
Plain Gemini 3.8 Live, which Google is not positioning as a flagship, scored 76.0 on the same leaderboard. That trails GPT-Live-1's Astra tier and its cheaper Sol tier, which scored 80.1. It does beat Google's prior live model, Gemini 3.1 Flash Live, which topped out at 71.5 on its high setting.
- Gemini 3.8 Live Ext. Thinking82.6 index points
- OpenAI GPT-Live-1 (Astra)81.5 index points
- Grok Voice Think Fast 2.081.3 index points
- Gemini 3.8 Live76 index points
- Gemini 3.1 Flash Live71.5 index points
Source: Artificial Analysis, accessed 2026-09-15
Google's own post cited two of its own benchmarks for Extended Thinking: 97.7% on Big Bench Audio, and 68.6% on a measure it calls agentic task completion. It published no comparable scores for plain 3.8 Live, and did not explain why the gap between its two new models is more than six points on Artificial Analysis' index.
What Google left out
Google's announcement carries no price for either model, unlike OpenAI, which priced GPT-Live-1 in the API at 5 cents a minute when it launched in September. Google also did not say which Gemini Enterprise customers can reach the private preview, or when a general release might follow. Its own post ties the release to the Gemini 3.8 Flash Cyber model it shipped two weeks earlier, but does not say whether the live-voice line shares any architecture with that text-based 3.8 family.
Sources
More in A.I.
- 01How a Heap Overflow and an SSO Bug Reached OpenAI's MonorepoHacktron chained a libheif image bug through OpenAI's own forum to hijack an engineer's Codex session and open a pull request in the internal openai/openai repository.
- 02Agility's Digit 5 Drops the Safety Cage, Not the SkepticismThe humanoid robot lifts 50 pounds and charges in 9 minutes, backed by $300 million in orders. An independent robotics writer says its business case still assumes a drop-in worker replacement.
- 03PrismML Shrinks a 27B Model to 5.9GB at 1.72 BitsTernary Bonsai 2 27B keeps 98.2% of its full-precision score by rebuilding Qwen3.8-27B's weights as three values instead of sixteen bits, and an independent tracker puts the retention slightly lower.
- 04OpenAI Discloses a Model That Wrote Its Own JailbreakAn unreleased Astra-family model added a fabricated persona to 27 training summaries this summer, and the successor model mostly ignored what it had written.