Releases

Aleph Alpha Releases Kolibri-1, a 78B Open-Weight MoE Model

Pastel vector hummingbird representing Aleph Alpha’s Kolibri-1 open-weight MoE model.
Kolibri-1 is a 78.1B open-weight MoE model focused on English and German.

Aleph Alpha has released Kolibri-1, an open-weight Mixture-of-Experts model for English and German. It has 78.1B total parameters, but only 3.46B of them are active for each token. For a model focused on just two languages, that is still a huge setup.

The weights are available on Hugging Face under the Apache 2.0 license, meaning anyone can download, run, and modify the model. Aleph Alpha released it on October 3, the Day of German Unity.

On paper, 3.46B active parameters sounds like something any decent GPU could handle. But unfortunately, real life is not that simple.

Mixture-of-Experts

A normal dense model uses all of its parameters for every token. A Mixture-of-Experts model, on the other hand, has a better way to handle the work with a limited number of active parameters. It has many smaller expert groups, and a router chooses only a few of them for each token.

Think of a company with hundreds of workers. Only a small group handles each task, so they can finish the job faster, while the company still keeps everyone on staff.

Kolibri-1 has 384 routed experts, uses 6 of them for each token, and always runs 1 shared expert. That is how a 78.1B model can do roughly the work of a much smaller 3.46B model for each token.

Kolibri-1 also supports adjustable reasoning effort, so you can choose how much thinking the model does for each request.

How Kolibri-1 Handles Long Context

Kolibri-1 can stretch all the way to 1,048,576 tokens, but that number needs some context of its own. Aleph Alpha trained the model up to 262,144 tokens, and that is still the safer range for harder tasks.

The model also avoids using full attention everywhere. Out of 50 layers, only 10 look across the full context. The other 40 use a 512-token sliding window. This keeps the KV cache from exploding as quickly, because most layers only need to remember a small recent chunk instead of the entire conversation. If KV cache is still a blurry term, I explained it in why local models need so much VRAM.

Can You Run Kolibri-1 Locally?

Yes, but not on a normal gaming PC. The 3.46B active parameter count makes Kolibri-1 sound lightweight, but that only describes how much of the model is used for each token. The full 78.1B weights still need to live somewhere, which is why the official FP8 checkpoint lands at around 78 GB.

Aleph Alpha mainly points to data center hardware such as H100, H200, and B200 GPUs through vLLM. That already tells you this is not exactly a model aimed at an average gaming PC.

There is a Q4_K_M GGUF at roughly 47.5 GB, while an MLX 8-bit build sits around 84 GB and targets Macs with 128 GB of unified memory.

Those builds are not official, though. Aleph Alpha has not tested them, and the MLX page still treats memory fit and speed as unverified.

The 47.5 GB GGUF is also a pretty good real-world example of quantization. With 78.1B parameters packed into that size, you end up at roughly 4.9 bits per weight, which lines up with what you would expect from Q4_K_M.

Because Kolibri-1 is an MoE model and only activates a small part of itself for each token, it may also tolerate system RAM offloading better than a dense model of similar total size. Still, I have not tested that setup myself, so I would treat it as an experiment rather than something guaranteed to work well.

Good Benchmarks, With Caveats

Aleph Alpha’s own tests put Kolibri-1 at 75.5 overall in English and 70.8 in German, ahead of models like Qwen3.6 35B-A3B, Nemotron 3 Super, and Mistral Small 4 in the same benchmarks. Reasoning also looks strong, with 96.9 on AIME 2025 and 84.3 on GPQA Diamond.

Tool calling is one of its weaker areas, though. Kolibri-1 scores 61.4 on BFCL v4 compared to 70.5 for Qwen3.5, which is worth knowing if you plan to use it heavily with tool calling. Also keep in mind that these numbers come from Aleph Alpha’s own testing.

Independent coverage also points out that some of the models it was compared against are already around six months old, and that the current open-weight leaders are still ahead.

Benchmarks are not really the whole point of Kolibri-1 anyway. Aleph Alpha built it around European use cases, with training infrastructure in Germany and Finland and a focus on areas like government, industry, and aerospace.

German is also a big focus. Aleph Alpha says its tokenizer uses around 11% fewer tokens than GPT-5’s tokenizer on German web text. The model is also trained to avoid answering when retrieved information does not support an answer, which can be useful for RAG.

There is no documented hosted API yet, so for now Kolibri-1 is mainly a self-hosted model through vLLM. BF16 weights are also available if you want to fine-tune it.

Sources

Aleph Alpha: Kolibri Has Landed

Kolibri-1 on Hugging Face

Aleph Alpha Technical Report

Updated Oct 4, 2026