
Aleph Alpha's Kolibri: A Documented Apache 2.0 Release
Aleph Alpha's German-English mixture-of-experts model ships with open weights and a detailed model card. The benchmark numbers are the company's own, and the sovereignty argument is a procurement case rather than a quality one.
Kolibri is a credible open-weight German-English model, but its quality claims remain unverified. Aleph Alpha's 78-billion-parameter mixture-of-experts model activates about 3.46 billion parameters per token and ships under Apache 2.0. Its benchmark scores are the company's own, and we found no independent reproduction.
The Weights Desk · 4 min read- Kolibri has 78.1B total parameters and about 3.46B active per token, per the Hugging Face model card.
- The Apache 2.0 license and the published model card make licensing, architecture and training compute checkable.
- Every benchmark figure, such as 96.9 on AIME 2025 and 66.4 on SWE-Bench, comes from Aleph Alpha; we found no independent reproduction.
- The company's two German-data figures differ: 21.3% of pre-training tokens in the blog post and about 23.9% in the model card mix; neither source reconciles them.
- The model card lists hallucinations, political bias and outdated knowledge among its limitations, so it suits human-supervised work rather than unsupervised deployment.
Aleph Alpha has released Kolibri, a German-English mixture-of-experts model with open weights under Apache 2.0, according to its Hugging Face model card and The Decoder's report. The release is well documented, but the performance claims are the company's own, and we found nothing showing a third party reproducing them. Our verdict: it is a serious option for supervised German-language work, and unproven everywhere else.
What is verifiable: architecture, license and compute
The structural facts are checkable because the model card publishes them. Kolibri has 78,103,074,560 total parameters and 3,457,573,120 active per token, across 50 layers with 384 experts per layer, one shared and six routed. Native context is 262,144 tokens, extendable to 1,048,576. Pre-training ran on 768 NVIDIA B200 GPUs for 21 days, with an estimated 950 MWh of energy including data-centre overhead, all per the model card.
What is not verified: the benchmark numbers
Every performance figure we found originates with Aleph Alpha. The model card lists 80.0 on MMLU-Pro, 96.9 on AIME 2025, 92.7 on HumanEval+ and 66.4 on SWE-Bench. The company's blog post reports a German overall score of 70.8 percent and says Kolibri competes with Qwen, Nemotron 3 Super and Mistral Small 4 models. We found no independent harness run, so these remain vendor claims and should be treated as leads, not settled results.
Where it breaks: the vendor's own caveats
The model card is direct about limits. Knowledge stops at June 18, 2026 in both English and German. The card lists systemic biases, political bias, hallucinations, outdated knowledge and the potential for harmful outputs despite safety training. Those are the failure modes a production team must test for, and the card offers no evidence of how often they occur. That gap, rather than any missing feature, is the main reason to hold off on unsupervised use.
Who should and should not use it
Kolibri fits teams that need German-language capability, weights they can host themselves and a permissive license. The model card lists a minimum of two 80 GB-class GPUs or a single H200, B200 or B300, so it is not a laptop model. Teams that need verified English-only frontier performance, independently confirmed agentic coding results or unsupervised high-stakes decisions have no evidence here to justify a switch. Run your own German evaluation set before committing.
What the sovereignty argument does and does not show
The sovereignty case is a procurement argument, not a quality one. Aleph Alpha's blog post says its teams built the model in Germany and trained it on infrastructure in Germany and Finland, under European and German law. Apache 2.0 weights give buyers auditable, self-hostable control that API-only models cannot. None of that tells a buyer how well the model performs. Watch for third-party German evaluations before treating sovereignty and quality as the same claim.
Verdict: pilot Kolibri where German output matters and a human reviews results, because the license, hardware profile and disclosed limits are concrete. Do not treat the launch scores as settled. Until an independent harness reproduces them, they are a lead, not a source.
- What is Kolibri and what hardware does it need?
- Kolibri is a German-English mixture-of-experts model with 78B total and about 3.46B active parameters per token, per its model card. The card lists a minimum of 2x A100 80 GB, 2x H100 SXM5, 1x H200, 1x B200 or 1x B300.
- Are the benchmark results independently verified?
- Not that we could find. The scores come from Aleph Alpha's model card and blog post. Treat them as leads until a third-party harness reproduces them.
- How German is the training data?
- Aleph Alpha's blog post says 21.3% of pre-training tokens are German, roughly 4.3T of 20T. The model card's data mix lists about 23.9% German. The origins do not explain the gap, which may reflect different stages or counting.
- Aleph Alpha releases Kolibri, an open-weight model that makes the case for European AI sovereignty — The Decoder
- Kolibri-1 model card — Aleph Alpha (Hugging Face)
- Kolibri has landed: a sovereign open-weight model — Aleph Alpha