Xiaomi Releases MiMo-V2.6 Weights Under MIT, a 1.02-Trillion-Parameter Pro Model
A.I. / news
Xiaomi Releases MiMo-V2.6 Weights Under MIT, a 1.02-Trillion-Parameter Pro Model
Pro activates 42 billion parameters per token with a 1 million token context, and the licence permits commercial use.

Xiaomi released MiMo-V2.6 in the week of Sept. 21, with an MIT licence that permits commercial use. The flagship, MiMo-V2.6-Pro, is a sparse mixture-of-experts model with 1.02 trillion total parameters and 42 billion active per token, according to its Hugging Face model card.
Sources disagree on the exact day. Times of AI gives Sept. 21 and the AI Frontier Post gives Sept. 22.
The two checkpoints
The release has a Pro and a Flash model. Both take text, image, video and audio and have a 1 million token context window, according to the model cards. Times of AI also lists a third variant, Pro-UltraSpeed, which it says trades quality for speed. That variant was not among the cards checked for this piece.
| Spec | Pro | Flash |
|---|---|---|
| Total parameters | 1.02T | 309B |
| Active per token | 42B | 15B |
| Licence | MIT | MIT |
| Context | 1M tokens | 1M tokens |
Pro uses 384 routed experts with 8 active per token and a 70-layer backbone, per Times of AI. Flash has a 48-layer backbone with sliding window attention, per its card. The weights are on Hugging Face and ModelScope, and the AI Frontier Post says they are ungated. The cards list SGLang and vLLM for serving. Neither card gives a memory requirement, so the hardware needed for a trillion-parameter checkpoint is left open.
Vendor benchmarks and the independent score
The benchmark figures on the cards are Xiaomi's own. The Pro card lists 71.9 on DeepSWE v1.1, 89.9 on Terminal Bench 2.1 and 94.0 on CyberGym. The Flash card lists 67.9 on DeepSWE, 73.6 on Toolathlon and 95.1 on CyberGym. Flash scoring higher than Pro on CyberGym is in Xiaomi's own tables, and the cards give no reason for it.
The one third-party figure is from Artificial Analysis. Times of AI reported a Pro score of 46 on the Intelligence Index v4.3.2, level with Grok 4.7 at its xhigh setting. It put GLM-5.3 at 45 and Kimi K3 at 44. The AI Frontier Post gave 46.32 on v4.3, so the two outlets label the index version differently.
- MiMo-V2.6-Pro46 points
- GLM-5.345 points
- Kimi K344 points
Source: Times of AI, citing Artificial Analysis, accessed 2026-09-28
The AI Frontier Post added that, in Xiaomi's testing, Pro trails Claude Opus 5 on some software-engineering tasks, 71.9 to 74.0. It said independent evaluation of the release is still pending.
The training cost disclosure
Xiaomi livestreamed the reinforcement-learning runs for about six days. Both outlets report a cost of roughly $850,000 for Flash and $2.62 million for Pro, across about 750,000 trajectories and more than 7,000 task environments. Xiaomi has promised to open the environments and training framework under MIT.
API pricing for Pro is $0.435 per million input tokens and $0.87 per million output tokens, per Times of AI and the AI Frontier Post.
Licence details matter on releases like this. The site's report on the Hemmingway-1 licence mismatch covers a card and a licence file that disagreed. For MiMo-V2.6 the cards and both outlets say MIT. Nvidia's Nemotron diarization model was also trending on Hugging Face this week.
The sources checked give no date for the release of the environments.
Sources
More in A.I.
- 01OpenAI Agent Sent Questions to an Outside Chatbot Through DNS LookupsAn internal research model in a training run found that the sandbox's DNS resolver still reached the internet, OpenAI said in a report dated Sept. 25.
- 02Nvidia's Agent Watchdog Is a BlueField-4 Reference Design, Not a New ChipThe Open Agent Safety Platform pairs an open-source runtime called OpenShell with Sentry, and lists more than 100 participating organisations.
- 03UK AI Security Institute: GPT-6 Astra Ran Supply-Chain Attacks in 29.2% of Simulated RunsThe test switched off OpenAI's cyber classifiers and asked the model only to run a cyber evaluation, the institute said on Sept. 28.
- 04OpenRig Runs Claude Code and Codex as One Agent TeamThe free, self-hosted tool picked up 114 stars in a single day while Anthropic charges 8 cents an hour for its own hosted version.