Ai2 Releases Olmo-core 3 With 2.7x MoE Training Speedup, Measured With Random Routing
A.I. / news
Ai2 Releases Olmo-core 3 With 2.7x MoE Training Speedup, Measured With Random Routing
The open training stack reports 52,000 tokens per second per GPU on a 47B-parameter mixture of experts. The 1.2-trillion-parameter result is a benchmark configuration, not a trained model.

The Allen Institute for AI (Ai2) released Olmo-core 3 on Oct. 1, a rebuilt open training framework for mixture-of-experts (MoE) language models, and reported a 2.7x throughput gain over its earlier approach on NVIDIA B300 GPUs.
The release, described in a post on the Hugging Face blog, includes code, a technical report and an interactive demo. Ai2 said it is the base for its next Olmo model, which will use an MoE design.

The 52,000 versus 19,400 tokens-per-second result
Ai2 switched from fully sharded data parallelism (FSDP) to distributed data parallelism, keeping experts resident on each GPU. It added expert parallelism, pipeline parallelism, GPU-resident routing and grouped GEMM kernels.
On eight B300 GPUs, a 47-billion-parameter MoE reached 52,000 tokens per second per GPU, against 19,400 for the previous FSDP implementation, according to Ai2's post. Unite.ai, which also read the report, labelled it a preliminary test.
- Earlier FSDP implementation19K tokens per second per GPU
- Olmo-core 352K tokens per second per GPU
Source: Ai2 post on the Hugging Face blog (vendor-supplied), accessed 2026-10-01
All figures in the post are Ai2's own. No outside group is named as having reproduced them.
Scaling from 8 experts to 128
Ai2 grew the expert pool from 8 to 128, with 4 experts chosen per token. Active parameters stayed near 3.2 billion while total parameters rose from 4.6 billion to 47 billion, and throughput fell by less than 5 percent.
The MXFP8 low-precision format raised throughput about 21 percent over a BF16 baseline and cut peak active memory from 103 GiB to 95 GiB. Activation recompute reduced peak memory by roughly a quarter at a cost of about a fifth of throughput.
The trillion-parameter figures are not trained models
Ai2 benchmarked a 1.2-trillion-parameter configuration across 512 GPUs at 858 TFLOP/s per GPU. It also ran a 2.38-trillion-parameter test with DeepEP v2.
Ai2 wrote that "these tests used random routing to measure system performance rather than the quality of a trained model." It described the 2.38-trillion run as a short capacity test rather than sustained training.
| Configuration | Result | What it is |
|---|---|---|
| 47B MoE, 8 B300 GPUs | 52,000 tokens/s per GPU | Preliminary throughput test |
| 1.2T MoE, 512 GPUs | 858 TFLOP/s per GPU | Benchmark, random routing |
| 2.38T MoE, DeepEP v2 | Not reported as throughput | Short capacity test |
No trained model at the 1.2-trillion or 2.38-trillion scale accompanied the release.
Licence and what to check before using it
The Olmo-core repository carries the Apache 2.0 licence and showed 1.6k stars and 321 forks when read. Its README says MoE support relies on the optional grouped_gemm dependency, which may need compiling from source until a specific pull request is released.
It also says published Docker images contain dependencies only, not the package. Unite.ai named Tianhua Tao as lead developer on a technical report with authors from Ai2 and the University of Washington.
This is a different kind of open release from weights such as PrismML's Ternary Bonsai 2 27B. Training code lets others reproduce a recipe, and Ai2 measured these numbers on B300 GPUs. For agent builders, the more immediate open release is NVIDIA's OpenShell sandbox.
Ai2 has not said when the next Olmo model will ship.
Sources
More in A.I.
- 01Qwen3.8-27B Ships Under Apache 2.0 and Fits in 17GB, but Spends 160 Million Tokens Where the Median Spends 43 MillionAlibaba's open-weight model scores 52 on Artificial Analysis's Intelligence Index. Its own benchmark figures are vendor-supplied, and users report slow runs.
- 02OpenAI Ties Moonshot AI to a July Campaign That Replayed Encrypted Reasoning, Offers No Evidence PubliclyOpenAI says 16,000 requests from more than 4,000 accounts tried to recover hidden model reasoning. Its attribution to Moonshot rests on its own assertion.
- 03Gemini 4 Argon Leads 13 of 18 Benchmarks Google Chose, but Only Cyber Defenders Can Use ItGoogle priced the model at $2 and $10 per million tokens and gave access first to its Fairwind Program, with a guardrail-free version for trusted defenders.
- 04Amazon Releases Strands Decider 2B, an Apache 2.0 Decision Model Built on Qwen3.5-2BAWS's Strands Labs scores 72.3 percent on JevBench at a 106 ms median on an RTX 3090. TypeSafe's CEO calls the current crop of rivals less serious than his own team.