Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in German
A.I. / news
Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in German
Aleph Alpha's 78B-parameter mixture-of-experts model activates 3.46B parameters per token, and its own model card shows a larger dense Qwen ahead on every headline benchmark.

Aleph Alpha released Kolibri-1 on Oct. 3, an Apache 2.0 mixture-of-experts model with 78 billion parameters, of which 3.46 billion are active for each token. Its own model card shows it scoring 70.8 on the company's German test mix, against 79.9 for Alibaba's dense Qwen3.8 27B.
A mixture-of-experts design routes each token to a subset of the model's parameters instead of all of them, which is why a 78B model can run at the cost of a roughly 3.5B one. The weights are downloadable on Hugging Face with no access form mentioned on the card. They ship in FP8 at about 78 GB, with a BF16 repository alongside.
What the Apache 2.0 licence covers
The card is specific about the limit of the licence. "The model weights are published by Aleph Alpha GmbH under Apache 2.0 license," it says, and then: "The license especially does not extend to underlying code, model architecture, parameter settings or any training method."
That matters in practice, because Kolibri-1 does not run on stock vLLM. The card says serving needs the aleph-alpha-inference package, which provides a Kolibri plugin for vLLM and pins the vLLM version it supports. Aleph Alpha also publishes a container image at ghcr.io/aleph-alpha/aleph-alpha-inference.
Hardware and context window
The card lists a minimum of two A100 80GB cards, two H100 SXM5 cards, or one H200, B200 or B300. It recommends two H100s, two H200s, one B200 or one B300. Each of its 50 layers holds 384 routed experts, six selected per token.
The context window is 1,048,576 tokens, but the card says native training length was 262,144 and recommends staying at or below that for serving efficiency. DataNorth, a Dutch marketing site that covered the release on Oct. 5, put it as "Treat the larger number as a ceiling, not a working size."
Training used about 20 trillion tokens: 62.5 percent English, 23.9 percent German and 13.6 percent code, with a knowledge cutoff of June 18, 2026. The model handles German and English only, and supports reasoning at configurable effort levels and tool calling.
How it scores against Qwen
Every figure below was run by Aleph Alpha at its highest reasoning setting, using its own evaluation framework. No third party has rerun them. The comparison column is the best other model in each row of the card's table.
| Benchmark | Kolibri-1 | Best other model in row |
|---|---|---|
| Overall, German | 70.8 | Qwen3.8 27B: 79.9 |
| Overall, English | 75.5 | Qwen3.8 27B: 80.2 |
| GPQA Diamond | 84.3 | Qwen3.8 27B: 89.2 |
| SWE-Bench Verified | 66.4 | Qwen3.6 35B-A3B: 73.8 |
| Terminal-Bench 2.1 | 27.7 | Qwen3.8 27B: 39.7 |
| Humanity's Last Exam | 21.5 | Qwen3.8 27B: 35.6 |
Against the smaller Qwen3.5 35B-A3B, DataNorth reports Kolibri-1 ahead on overall German (70.8 against 69.8) and overall English (75.5 against 74.7), and behind on German company-document retrieval (67.5 against 70.0).
- Qwen3.8 27B (dense)79.9 points
- Kolibri-1 (78B MoE)70.8 points
- Qwen3.5 35B-A3B69.8 points
Source: Aleph Alpha Kolibri-1 model card and DataNorth, accessed 2026-10-08. Vendor-run, not independently reproduced.
The terminal-agent row is the widest relative gap: 27.7 for Kolibri-1 against 39.7 for both Qwen3.5 35B-A3B and Qwen3.8 27B, according to the card. DataNorth's closing assessment was that "if quality matters more than hardware, Kolibri is not the strongest open option."

Where it fits among open-weight releases
The case for Kolibri-1 is provenance, not rank. The weights come from a German company, the training data mix was disclosed, and DataNorth reports that Aleph Alpha has signed the EU's Code of Practice for general-purpose A.I. and published a data summary.
For teams that only need the best open model that fits on a single card, Qwen3.8-27B is also Apache 2.0, has a 262,144-token window and already had 482 finetunes when The Terminal covered it. It needs fewer parameters in memory and scored higher in every row of the table above. The card lists no hosted API or pricing for Kolibri-1, so there is no direct comparison with a closed model such as Claude Haiku 5.5, which Anthropic prices at $0.10 per million input tokens.
The next number to watch is independent: whether anyone outside Aleph Alpha posts a German-language result for Kolibri-1 that lands within a point of the card's 70.8.
Sources
More in A.I.
- 01LTX-2.5's Free Commercial Licence Stops at $10 Million in Group Revenue, and Its Card Asks for Contact DetailsLightricks' open-weight video and audio model is free for production use below that line, but the revenue test counts parent companies and affiliates, and the card publishes no benchmark scores.
- 02Gemini 4 Argon Costs $1.99 a Task at Promo Price Against $0.72 for GPT-6.1 Sol, and Is Not Yet on SaleGoogle's introductory $2 and $10 rates match OpenAI's Sol per token, but Artificial Analysis figures cited by eesel show Argon writing 62,000 output tokens a task where GPT-6 Astra writes 27,000.
- 03EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 ProGoogle DeepMind's Apache 2.0 embedding model has 740M parameters in three modular pieces, and the only benchmark number its launch post prints is a 9.92-point gain on MTEB Code.
- 04Qwen3.8-27B Is Apache 2.0 With a 262,144-Token Window and Already Has 482 Finetunes, Including Cloudflare's ClefAlibaba's dense 27B model claims 61.7 on SWE-bench Pro against 53.4 for Opus 4.6 Max, and the model card says nothing about training data.