Aleph Alpha releases Kolibri open-weight German model
The German-English mixture-of-experts model carries 78 billion parameters with about three billion active per token, released under Apache 2.0.
In short
Aleph Alpha has put Kolibri into the open: a German-English mixture-of-experts model whose 78 billion parameters narrow to roughly three billion active per token, with the weights downloadable under Apache 2.0.
At a glance
- 78 billion parameters in total, about three billion active per token through a mixture-of-experts router.
- German makes up 21.3 percent of the training data, fed by a German data pipeline built for the project.
- Trained on 768 B200 GPUs in Germany and Finland, with context windows of up to one million tokens.
- Scores 71 percent on German benchmarks and decodes faster than GPT OSS A5B, Qwen 3.6 A3B and Gemma 4 A4B.
- Weights ship under Apache 2.0 on Hugging Face as Aleph-Alpha/Kolibri-1.
Aleph Alpha has released Kolibri, a German-English mixture-of-experts model with 78 billion parameters of which only around three billion fire per token. The weights are on Hugging Face as Aleph-Alpha/Kolibri-1 under Apache 2.0. For a European lab, that combination of scale and permissive licensing is the news, not the benchmark line.
Capacity without the inference bill
Sparse routing is the point here: the model holds 78 billion parameters but spends roughly three billion per token, which keeps serving cost closer to a small dense model than to its parameter count. Context windows reach one million tokens. Aleph Alpha reports 71 percent on German benchmarks with faster decoding than GPT OSS A5B, Qwen 3.6 A3B and Gemma 4 A4B.
A model built around German, not translated into it
German accounts for 21.3 percent of the training data, assembled through a dedicated German data pipeline. Chinese models were used to generate part of the synthetic training data, which complicates the clean sovereignty story the release is framed around. Training ran on 768 B200 GPUs across sites in Germany and Finland.
Who is meant to run it
The stated targets are public administration, aviation and industry. Kolibri was developed under European law with the EU AI Act in view, and the Apache 2.0 licence is what makes that framing operational: a ministry or an airline can host the model itself instead of routing records through someone else's API.
What stays unverified
Several details cannot be checked against the one source that was readable. The number of experts, the training token count, the tokenizer, and the task mix behind that 71 percent figure are all unstated. A second outlet returned HTTP 403 and could not be read, so nothing from it is cited here. Pricing, who paid for the GPU time, and whether base, instruct or reasoning variants shipped separately remain unconfirmed.
Figures in this article come from The Decoder.
FAQ
What is Aleph Alpha Kolibri?
A German-English mixture-of-experts model with 78 billion total parameters, about three billion of them active per token, and context windows of up to one million tokens.
Where can I download the Kolibri weights?
On Hugging Face as Aleph-Alpha/Kolibri-1, under the Apache 2.0 licence, which permits self-hosting inside your own data center.
How well does Kolibri perform in German?
Aleph Alpha reports 71 percent on German benchmarks with faster decoding than GPT OSS A5B, Qwen 3.6 A3B and Gemma 4 A4B. The underlying task list is not documented in the available source.