Aleph Alpha Kolibri: Open-Weight German-English Model

Aleph Alpha Kolibri: Open-Weight German-English Model

German AI company Aleph Alpha has published Kolibri, a bilingual German-English language model with openly available weights. The company is pitching it as a practical tool for public administration, aviation and industry. It also sees the model as evidence that Europe can build competitive AI on its own terms.

Kolibri is the German word for hummingbird. According to the company's technical report, the model was trained on European hardware, developed under European law and designed with the EU AI Act in mind. The AI Act is the European Union's framework for regulating AI systems.

A large model that runs light

On paper, Kolibri has 78 billion parameters. In practice, it uses only about three billion of them for any single token. This comes from a mixture-of-experts architecture, which splits the network into specialised sub-networks and sends each token through only a few of them.

The trade-off is simple. A model built this way can hold a lot of knowledge in its total parameter count while keeping the compute cost per token closer to that of a much smaller model. That approach has become common across the industry, and tools for training mixture-of-experts models more efficiently are getting attention of their own.

Other details from the release:

  • Context window: up to one million tokens.
  • Training hardware: 768 Nvidia B200 GPUs located in Germany and Finland.
  • License: Apache 2.0, a permissive open-source license that allows commercial use.
  • Distribution: weights are available on Hugging Face.

The quality-versus-cost claim

Aleph Alpha's central performance claim is that Kolibri sits on the Pareto front for both German and English. In plain terms, the company says that among the models it compared, none delivers better quality at the same operating cost, and none runs cheaper at the same quality.

The company says Kolibri beats comparable models with similar architectures on either quality or cost. Some of those comparison models are significantly older, with release dates in March and April 2026. That detail matters when reading the claim. In a field that moves this fast, a gap of several months can shift the benchmark picture.

Built for German

The bilingual focus is not just marketing. German makes up 21.3 percent of the training data. Aleph Alpha built a dedicated German data pipeline for the project to source and prepare that material.

Most frontier models are trained mainly on English text, and their performance in other languages often trails behind. For German government offices or manufacturers that work largely in German, a model with a deliberate German foundation could be a better fit than a general-purpose system with German added on top.

A sovereignty model with Chinese inputs

There is one detail that complicates the story. Aleph Alpha used Chinese models to generate synthetic training data for Kolibri. Synthetic data, meaning text produced by other AI systems and used to train new ones, is now a standard technique. Still, it sits awkwardly next to a sovereignty message.

It stands out all the more because Aleph Alpha itself has published research on how Chinese models echo state doctrine. The release does not say how the company filtered or checked the synthetic data, so readers should treat that as an open question.

Our Take

Kolibri is a clear signal about where Aleph Alpha wants to compete. It is not trying to match the largest US labs on raw capability. It is selling efficiency, language fit, legal alignment with EU rules and the reassurance of European-trained weights. For organisations in regulated sectors such as public administration and aviation, that combination may matter more than topping a general leaderboard.

The architecture also fits a broader trend. Sparse models with a few billion active parameters are increasingly popular because they are cheaper to run, as seen with releases like Qwen's compact coding model. The Apache 2.0 license lowers the barrier further, since companies can deploy the model on their own infrastructure without negotiating terms.

Some caution is warranted. The performance claims come from Aleph Alpha's own report, and the comparison set includes models that are several months old. It is worth watching whether independent benchmarks confirm the Pareto-front claim against newer competitors. Another open question is whether public bodies and industrial firms actually adopt Kolibri in production. The use of Chinese-generated synthetic data may also draw questions from the very customers who care most about sovereignty. How Aleph Alpha answers them could shape how credible the European AI argument looks.