Aleph Alpha Releases Kolibri, a 78B German-English Model, Under Apache 2.0
A.I. / news
Aleph Alpha Releases Kolibri, a 78B German-English Model, Under Apache 2.0
Only 3.46B of the 78.1B parameters are active per token, and the FP8 weights need about 78 GB of memory.

Aleph Alpha released Kolibri, an English-German mixture-of-experts model, on Oct. 3 under the Apache 2.0 licence, with the weights on Hugging Face. The model has 78.1 billion parameters in total and activates 3.46 billion per token.
The Heidelberg company said in its launch post that it trained and developed Kolibri on infrastructure in Germany and Finland. The post lists Ilhan Scheer as chief executive. The Hugging Face card says the licence covers weights and configurations only, and that Aleph Alpha keeps the rights to its architecture and methods.
Kolibri hardware: 78 GB of FP8 weights
The weights ship in FP8, which puts the memory footprint at about 78 GB. The model card lists the minimum as two A100 80 GB, two H100 SXM5, one H200, one B200 or one B300.
Native context is 262,144 tokens. The card says it extends to 1,048,576. The model has 50 layers, each with 384 experts of which 6 are routed per token, plus one shared expert. The card describes three reasoning-effort settings and Hermes-style function calling. Its knowledge cutoff is June 18, 2026.
| Spec | Kolibri-1 |
|---|---|
| Parameters | 78.1B total, 3.46B active |
| Context | 262,144 native, 1,048,576 extended |
| Weights | FP8, about 78 GB |
| Licence | Apache 2.0 |
Training data: 20 trillion tokens, 21.3% German
Aleph Alpha filtered more than 200 trillion raw tokens down to 20 trillion for pre-training, then added 3.44 trillion mid-training tokens and about 200 billion long-context tokens. German made up 21.3% of the pre-training mix, roughly 4.3 trillion tokens. English was about 62% and code about 14%.
Pre-training took 21 days on 768 NVIDIA B200 GPUs, or 392,000 GPU hours, according to the card. The company did not publish training cost or the list of data sources.
Kolibri benchmarks are Aleph Alpha's own
All scores below come from Aleph Alpha and were not run by an independent party. The company reports 96.9% on AIME 2025 in English and 87.5% in German, 84.3% on GPQA Diamond and 92.7% on HumanEval+. On AA-Omniscience it reports a 44.0% non-hallucination rate. The card gives overall scores of 75.5 for English and 70.8 for German.
RuntimeWire counted the German Basic Law at 35,190 tokens with Kolibri's 128,000-entry tokenizer, against 41,482 for the tokenizers behind GPT-4o and GPT-5.
- Kolibri tokenizer35K tokens
- GPT-4o / GPT-5 tokenizers41K tokens
Source: RuntimeWire, accessed 2026-10-03
The card says the model is meant for assistants and agent workflows with human review, not autonomous operation.
Where Kolibri sits among open weights
Open-weight releases on this site have mostly been small or narrow, such as Cloudflare's 27B Clef decision model and Runway's Praxis-1. Kolibri is a general model at a size that needs a data-centre GPU, which is the point for the regulated customers Aleph Alpha is aiming at.
Aleph Alpha has published its own comparison with similar-sized models, and no independent leaderboard result for Kolibri had appeared as of Oct. 3.
Sources
More in A.I.
- 01GPT-6.1 Sol Is Priced at One-Fifth of Astra, and Its System Card Rates Cyber CriticalOpenAI's Sept. 29 addendum also shows the model misrepresenting its own coding work more often than GPT-6 Astra did.
- 02Gemini 4 Argon Goes to Cyber Defenders First, With Broad Access UndatedGoogle priced its new frontier model at $2 and $10 per million tokens for an introductory period, then $4 and $20, and has not said when most developers get it.
- 03Aleph Alpha's Kolibri Ships Under Apache 2.0, Compared Only With Spring ModelsThe 78B-parameter German-English model activates 3.46B per token, and its published benchmark table leaves out every open-weight release since the spring.
- 04Runway's Praxis-1 Robot Model Is Open-Weight on Paper, With Weights Still UnreleasedThe video-trained control model is being tested by Noble Machines, Standard Bots and Ultra, and Runway has not published a parameter count or a licence.