Aleph Alpha's Kolibri-1 Is a 78B Apache 2.0 Model That Activates 3.46B Parameters, and Its Training Data Stays Closed
A.I. / news
Aleph Alpha's Kolibri-1 Is a 78B Apache 2.0 Model That Activates 3.46B Parameters, and Its Training Data Stays Closed
The German lab published weights on October 3, but its blog and its model card disagree on active parameters and on how much of the corpus is German.

Aleph Alpha released Kolibri-1 on October 3, a 78-billion-parameter mixture-of-experts model for German and English, with weights on Hugging Face under the Apache 2.0 licence. The model card lists 3.46 billion parameters active per token and a native context window of 262,144 tokens.
The weights are not gated behind a form. The training data is not published, and Aleph Alpha said on its launch post that its fine-tuning datasets remain proprietary.

What Kolibri-1 is, in numbers
The card describes a 50-layer model with 384 experts per layer and a 4:1 mix of sliding-window and grouped-query attention. Aleph Alpha's launch post says six experts are active per token.
| Spec | Kolibri-1 | Source |
|---|---|---|
| Total parameters | 78B | Model card |
| Active per token | 3.46B | Model card |
| Native context | 262,144 tokens | Model card |
| Extended context | 1,048,576 tokens | Model card, by extrapolation |
| Pre-training tokens | 20T | Model card |
| Licence | Apache 2.0 | Model card |
The card adds 3.44 trillion tokens of mid-training and 201 billion tokens of long-context training. It lists minimum hardware of two A100 80GB cards, two H100 SXM5 cards, or one H200, B200 or B300.
Where the two documents disagree
The launch post and the card do not describe the same model in the same numbers. The blog rounds active parameters to 3 billion, where the card says 3.46 billion.
The corpus split differs as well. The card gives 62.5% English, 23.9% German and 13.6% code. The blog gives 21.3% German, about 62% English and about 14% code, which sums to 97.3%.
Aleph Alpha did not say which figure is current. The card is the document attached to the weights, so this report uses it for the table above.
Training run and benchmark claims
Aleph Alpha said it trained Kolibri-1 on 768 Nvidia B200 GPUs in Germany and Finland. Pre-training took 21 days, and the blog said 38 unplanned interruptions in that period were handled automatically.
The model card reports 80.0 on MMLU-Pro with chain of thought in English, 96.0 on AIME 2026, 92.7 on HumanEval+ and 66.4 on SWE-Bench Verified. All four are vendor-supplied, and the card names no outside party that reran them.
The Decoder, in a report by Matthias Bastian dated October 5, said Kolibri scores 71% on German benchmarks. It also said the model decodes faster than GPT OSS A5B, Qwen 3.6 A3B and Gemma 4 A4B. Both claims trace to Aleph Alpha.
Synthetic data from Chinese models
The Decoder reported that Chinese models generated part of the synthetic training data. Aleph Alpha markets Kolibri for what it calls sovereign use in public administration, aviation and industry, and the question of which models wrote that data matters to buyers with sourcing rules.
The launch post did not name the models or the share of the data they produced. The company said Kolibri was built "with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind from the ground up".
How it sits beside other open releases
Kolibri-1 is smaller in active parameters than Reflection's Beam, which has 501B parameters and 23B active and whose weights had not shipped as of that report. Kolibri's weights are downloadable now.
That contrasts with Runway's Praxis-1, a robot model announced without weights. Of the three, only Kolibri-1 can be run today.
Aleph Alpha said enterprise deployment and specialisation are available through its sales contact, and that the model runs on-premise through its aleph-alpha-inference package with a vLLM plugin. The launch post gave no date for any further release.
Sources
More in A.I.
- 01Reflection Announces Beam, a 501B Open-Weight Model, but Has Not Named a LicenceThe weights are promised for later in October. Reflection's own post gives scores and training scale, and says nothing on licence terms or API pricing.
- 02Meta Disputes Inc. Columnist's Claim That Muse Read His Messages While Mac Researcher Calls It a BackdoorTwo separate Muse problems are being reported as one. A researcher's local-access flaw is documented; the claim that the agent read private messages is contested by Meta.
- 03Anthropic Staff Reported a Claude Chat to Police, and a Florida Woman Faces a Felony ChargeThe company's privacy policy allows disclosure to prevent serious harm. Its transparency report counted zero emergency requests from police, and does not count referrals it makes itself.
- 04Vals AI's 90 Claude Opus 5.5 Agents Name Two Magnetic Semiconductor Candidates, Neither Yet MeasuredOne candidate was first synthesised in 1999 and has a measured ordering temperature of 376 K. The other may not survive the furnace.