Aleph Alpha's Kolibri Ships Under Apache 2.0, Compared Only With Spring Models
A.I. / news
Aleph Alpha's Kolibri Ships Under Apache 2.0, Compared Only With Spring Models
The 78B-parameter German-English model activates 3.46B per token, and its published benchmark table leaves out every open-weight release since the spring.

Aleph Alpha released Kolibri on Oct. 3, a 78.1 billion-parameter English and German model whose weights anyone can download under the Apache 2.0 licence. The company published its own benchmark table against three open-weight rivals, all of which date from the spring.
The Aleph Alpha announcement describes a mixture-of-experts design, which routes each token to a few of the model's parameter groups instead of all of them. The model card on Hugging Face shows the repository is not gated.
What the model card specifies
Kolibri has 384 experts with six active per token, which means 3.46 billion of its 78.1 billion parameters run on any one token. Native context is 262,144 tokens, and Aleph Alpha says it is validated to 1 million.
The weights are stored in 8-bit floating point and total about 78GB. The card lists the minimum hardware as two A100 80GB cards, two H100 SXM5 cards, one H200, one B200 or one B300. Aleph Alpha says two H100s serve 18 concurrent requests of 256,000 tokens each.

Pre-training used 20 trillion tokens on 768 B200 GPUs in Germany and Finland. Mid-training added 3.44 trillion tokens and long-context work added about 200 billion. The knowledge cutoff is June 18, 2026.
The company did not publish its training data sources. It describes how the data was curated but names none of it.
The two documents also disagree. The announcement puts the pre-training mix at 62 percent English, 21.3 percent German and 14 percent code. The model card says 62.5 percent, 23.9 percent and 13.6 percent.
The benchmark table and who is missing
Every score below is Aleph Alpha's own. On AIME 2026, a maths competition set, it reports 96.0 for Kolibri against 91.0 for Alibaba's Qwen3.6-35B. On GPQA Diamond, a graduate science test, it reports 84.3 against 83.4.
The table is not uniformly favourable. On HumanEval+, a coding test, Kolibri scores 92.7 and Nvidia's Nemotron 3 Super scores 94.7. On the AA-Omniscience Index, which penalises wrong answers, Kolibri scores minus 32.8 and Qwen3.6-35B scores minus 15.3.
- Kolibri, AIME 202696 points
- Qwen3.6-35B, AIME 202691 points
- Kolibri, GPQA Diamond84.3 points
- Qwen3.6-35B, GPQA Diamond83.4 points
Source: Aleph Alpha announcement, Oct. 3, 2026. Vendor-supplied, not independently run.
Jakob Steinschaden of Trending Topics wrote that the comparison set is Qwen3.6-35B-A3B, Nemotron 3 Super 120B-A12B and Mistral Small 4 119B-A6B. "All three comparison models date from the spring," he wrote, adding that Chinese labs have since released stronger open-weight models, among them Qwen3.8, GLM-5.3, Kimi K3 and MiMo-V2.6-Pro.
Steinschaden estimated that a model at Qwen3.6 level would score roughly 15 to 20 on the Artificial Analysis index, behind at least 20 other open-weight models. He put the leader, MiMo-V2.6-Pro, at 46. That is his estimate, not an Artificial Analysis listing for Kolibri.
How it compares with what people download
The Hugging Face listing for Kolibri-1 shows 355 likes and 1,135 downloads in the last month. Alibaba's Qwen3.8-27B, also Apache 2.0 and ungated, shows 16,900 likes and more than 6.8 million.
Aleph Alpha is not selling to that audience. It calls Kolibri a model for sovereign, mission-critical work and says it was built and trained in Germany and Finland. It reports internal proxy scores for an automotive supplier task at 0.99, semiconductors at 0.80 and the German public sector at 0.75.
Those three tasks are internal and unpublished, so no outsider can rerun them. Aleph Alpha did not say which customers are using the model.
Aleph Alpha's post does not cite an independent evaluation of Kolibri. Cloudflare's Clef and Runway's Praxis-1 are two other open-weight releases this site has covered, for readers who want a comparison of how each lab documents its weights.
Sources
More in A.I.
- 01GPT-6.1 Sol Is Priced at One-Fifth of Astra, and Its System Card Rates Cyber CriticalOpenAI's Sept. 29 addendum also shows the model misrepresenting its own coding work more often than GPT-6 Astra did.
- 02Gemini 4 Argon Goes to Cyber Defenders First, With Broad Access UndatedGoogle priced its new frontier model at $2 and $10 per million tokens for an introductory period, then $4 and $20, and has not said when most developers get it.
- 03Runway's Praxis-1 Robot Model Is Open-Weight on Paper, With Weights Still UnreleasedThe video-trained control model is being tested by Noble Machines, Standard Bots and Ultra, and Runway has not published a parameter count or a licence.
- 04OpenAI Safety-Report Lead David Robinson Quits and Calls the Culture BrokenRobinson's Atlantic essay points to the July Hugging Face breach by about 700 OpenAI agents, and the company has answered with one spokesperson statement.