Nvidia's Nemotron 3 Cuts Speaker-ID Errors by 41%
A.I. / news
Nvidia's Nemotron 3 Cuts Speaker-ID Errors by 41%
The open-weight model doubles the speaker count of its predecessor but got slightly worse on one two-speaker test.

Nvidia released Nemotron 3 Diarization on Sept. 23, a 100-million-parameter open-weight model that identifies who is speaking in audio with up to eight speakers, under the OpenMDW License 1.1, which permits commercial use.
The model replaces Nvidia's Streaming Sortformer, which topped out at four speakers. Nvidia's own benchmarking found an average 41 percent relative reduction in diarization error rate against that predecessor across eight evaluation conditions at 1.04-second latency, with gains ranging from 9 percent on the CALLHOME-Part2 test to 65.2 percent on NOTSOFAR1 MHM. On the DIHARD III evaluation, the new model scored a 12.73 percent error rate at its slowest, most accurate setting, against 19.09 percent for the baseline.
Where the numbers came from, and where they slipped
MarkTechPost reported that Nemotron 3 ranked first on Voice Arena's Diarization-Bench, an independent test covering 139 English conversations and about 22 hours of audio, with a 14.72 percent error rate against 19.3 percent for the next-best system, a roughly 24 percent relative improvement. Not every result moved in the model's favor: Nvidia's own documentation shows the two-speaker CALLHOME test at 30.4-second latency got slightly worse, with the error rate rising from 5.68 percent to 5.98 percent against the predecessor.
| Latency setting | Error rate | Throughput (RTFx) |
|---|---|---|
| Offline, 30.4s buffer | 12.73% | 15,113x |
| Low, 1.04s | 13.18% | 865x |
| Ultra-low, 0.32s | 13.55% | 292x |
Nvidia trained the model on roughly 10,000 hours of real conversations combined with 82,611 hours of simulated multi-speaker audio spanning 21 languages. Licensed recordings from the company David AI cut the compound error rate further, from 11.19 percent to 10.42 percent, according to the model card.
The model card also lists a speaker-counting accuracy of 81.47 percent on the DIHARD III test, against 75.29 percent for the predecessor, meaning it more often gets the raw number of people in a recording right before it even tries to label who said what. Nvidia built in four latency presets, from an ultra-low 320-millisecond mode meant for live captioning down to the slower 30.4-second offline buffer used for the most accurate transcripts, so a developer can trade accuracy for speed depending on whether the audio is being processed live or after the fact.
What still trips it up
Nvidia's documentation states plainly that recordings with more than eight speakers can produce missed or misassigned speech, and that heavy background noise, reverberation and far-field microphone capture all raise the error rate. The company positions the model for meeting transcription tools, call analytics, podcast production pipelines and the kind of speaker memory a voice agent needs to track who it is talking to, rather than for open-ended crowd recordings.
The model runs through Nvidia's NeMo framework on Ampere, Ada Lovelace, Hopper or Blackwell GPUs; Nvidia tested it on an RTX PRO 5000 card at BF16 precision. It joins a run of open-weight audio releases this quarter: OpenAI's GPT-Live-1 API launched Sept. 10 at 5 cents a minute for real-time voice, and a licensing mismatch inside VoiceStudio's default voice-cloning model drew scrutiny earlier in September. Nemotron 3 answers a narrower question than either: not what a voice sounds like, but who is speaking and when.
Sources
More in A.I.
- 01OpenRig Runs Claude Code and Codex as One Agent TeamThe free, self-hosted tool picked up 114 stars in a single day while Anthropic charges 8 cents an hour for its own hosted version.
- 02OpenAI Says Agents Leaked 53 ChatGPT User ImagesThe company's Sept. 25 update says it still cannot match the images to the accounts that made them.
- 03Altworld's Hemmingway-1 Isn't Apache-Licensed, Despite ReportsHugging Face's own metadata says the 27-billion-parameter writing model is noncommercial only, contradicting at least one widely read AI blog.
- 04PrismML Shrinks Qwen3.8 27B to 5.9GB, Keeps 98% of Its ScoreBonsai 2 27B compresses Alibaba's Qwen3.8 27B to ternary weights averaging 1.76 bits each, and the Caltech-founded startup says it scores 83.9 against the original's 85.4 on its own benchmark suite.