Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Open-Weights Releases

Aleph Alpha open-sources Kolibri-1, a 78B sovereign EU model

Aleph Alpha released Kolibri-1, a 78B English-German open-weight model under Apache 2.0, with 3.46B active parameters per token and a 1M-token context.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A single rack server in a dim European datacenter aisle, one amber status LED lit among rows of unlit white indicators, for Aleph Alpha's Kolibri-1, a 78B open-weight English-German model released under Apache 2.0 and sized to run on two H100s or one H200 in a customer's own machine room. Illustration: DeAI
A single rack server in a dim European datacenter aisle, one amber status LED lit among rows of unlit white indicators, for Aleph Alpha's Kolibri-1, a 78B open-weight English-German model released under Apache 2.0 and sized to run on two H100s or one H200 in a customer's own machine room. Illustration: DeAI

Aleph Alpha, the Heidelberg-based AI company, released Kolibri-1 on October 3 with full weights on Hugging Face under the Apache 2.0 license: a 78B-total-parameter English-German mixture-of-experts model that activates only 3.46B parameters per token. It is the first fully downloadable open-weights model from a European lab aimed at regulated on-premises deployment, and its hardware floor starts at a single H200.

Key facts

What happened

The release landed on the Day of German Reunification, a date Aleph Alpha chose deliberately. The launch post frames Kolibri as "a sovereign open-weight model" for public administration, industrials, and aerospace: customers who run models in their own machine rooms and do not want data leaving their jurisdiction. Every load-bearing artifact was published at once — the full weights and model card on Hugging Face, a technical report, a public training-data summary in the European Commission's template, and the serving stack as a container image.

The architecture is a 50-layer mixture-of-experts transformer with 384 experts per layer, 6 routed plus 1 shared, using 4:1 sliding-window-to-grouped-query attention. The model card lists FP8 block-quantized weights with bfloat16 embeddings, norms, and router. Aleph Alpha shipped a BF16 sibling alongside the FP8 release, and community quantizations appeared within the first day. Serving runs through vLLM with a custom plugin in the aleph-alpha-inference package; the model card documents Hermes-style tool calling, a reasoning mode with low/medium/high effort settings, and an OpenAI-compatible endpoint.

The German-language engineering is the part with the least direct competition. Aleph Alpha built a bilingual tokenizer trained on organic German text rather than translations — 21.3% of pre-training tokens are German, with only 6% translated material — and the company's argument is that translated text carries the cultural fingerprint of its source language. It also trained the model to abstain, answering "I don't know" when a document does not contain the answer, which is the behavior regulated buyers in document-processing workflows consistently ask for.

One more thing worth knowing before the sovereignty framing: Aleph Alpha is not an independent company for much longer. It signed a definitive business combination agreement with Cohere on September 16, creating a dual-headquartered company operating as Cohere, with Toronto and Berlin as dual headquarters and Heidelberg retained as a research center. Kolibri is likely the last frontier-adjacent model released under the Aleph Alpha name, at least under independent ownership.

Why it matters

For builders deciding where to run models, Kolibri-1 fills a specific slot rather than resetting the leaderboard. The 3.46B active-parameter count puts it in the small-active class that serves cheaply at high volume; the ~78 GB FP8 footprint puts it on one H200 or two H100s, which is a single accelerator node rather than a cluster. Apache 2.0 puts it in the most permissive license class available — compare the MIT-licensed GLM-5.3-Flash or the open-weight versus open-source distinction that governs what a download actually grants. Nothing else in that class was trained specifically for German public-sector language.

Treat the benchmark table the way this publication treats every vendor benchmark: as Aleph Alpha's self-reported results. The launch post says Kolibri matches models with up to four times its active parameter count and sits on the Pareto frontier for quality versus serving cost in both languages. Those are the vendor's own claims, not independently verified measurements, and Kolibri was not yet listed on Artificial Analysis at release.

The comparison field also deserves scrutiny. Aleph Alpha benchmarked against Qwen3.6-35B-A3B, Nemotron 3 Super, and Mistral Small 4, all spring releases. Trending Topics' analysis points out what the table omits: the current open-weight leaders — Xiaomi's MiMo-V2.6-Pro, Z.ai's GLM-5.3, Moonshot AI's Kimi K3 — score in the 44-46 range on the Artificial Analysis Intelligence Index, where the comparison models sit at 11-18 points. Even within its own weight class, Qwen3.8-Flash-Next reaches 40 points with 6B active parameters. On public benchmarks, Kolibri is plausibly a mid-tier open model. Its bet is that German-first quality, abstention behavior, and deployment freedom matter more to its buyers than a leaderboard position.

The EU compliance angle is concrete, not decorative. Aleph Alpha is a signatory of the EU GPAI Code of Practice, and the release includes the Copyright compliance contact and the public training-data summary that the GPAI transparency obligations now require. A fully open release with complete documentation is the cheapest way to satisfy those obligations, and other European buyers shopping for sovereign options — including NVIDIA's Nemotron line in US sovereign deployments — now have a European-trained alternative with a permissive license.

Background

Aleph Alpha has been rebuilding its model program around a pipeline it calls the Model Factory. The launch post describes validating that pipeline with Kolibri Origin, a 30B-total/3B-active model with a 65,536-token context that was never publicly released, then running Kolibri through the same pipeline — hundreds of ablation experiments and a pre-training run that recovered from hardware failures without human intervention. Only about three months separate the end of Kolibri Origin's training from Kolibri's release.

The sovereign-AI angle is where this release lands on a beat DeAI follows closely. What are open-weight models covers why downloadable weights change the deployment calculus for regulated buyers, and IBM's self-hosted Bob release showed the same demand from the tooling side: enterprises want "bring the AI to the data" as a purchasable default. Kolibri is the first model trained specifically for that posture by a European vendor, on disclosed European infrastructure.

The merger changes the trajectory, though. Cohere has its own sovereign push and its own model line; the combined company's blog promises continuity for Aleph Alpha's government work, but roadmap continuity across an acquisition is a claim, not a commitment. Whether Apache 2.0 remains the licensing norm for Heidelberg's output after close is one of the concrete things to check.

What's next

Three checkpoints. First, independent benchmarks: Kolibri is downloadable and runnable, so Artificial Analysis listings, community eval runs, and German-language evaluations against Qwen3.8-class models are all possible within weeks, and they are the only thing that converts vendor claims into shared facts. Second, whether serving costs in practice match the Pareto-frontier framing — the FP8 footprint and 3.46B active parameters make claims testable against real throughput numbers. Third, the Cohere transaction: the deal is expected to close later this year, and Heidelberg's licensing and release practices after integration will say whether Kolibri was a capstone or a new pattern.

Questions

What is Kolibri-1?
Kolibri-1 is an open-weight English-German mixture-of-experts language model from Aleph Alpha, released October 3, 2026 under Apache 2.0. It has 78B total parameters with 3.46B active per token, supports up to 1,048,576 tokens of context, and ships full weights plus configuration on Hugging Face.
Is Kolibri-1 really open source?
The weights and configuration files are published under Apache 2.0, which allows commercial use, modification, and redistribution. The model card states the license does not extend to Aleph Alpha's training code, training methods, or other artifacts, so it is an open-weights release rather than fully open source.
What hardware does Kolibri-1 need to run?
Per the model card, the FP8 weights take about 78 GB. The minimum configuration is 2x A100 80GB or 2x H100, or a single H200, B200, or B300. Aleph Alpha recommends 2x H100 or 2x H200, or one B200/B300, and serves it through vLLM with a custom plugin.
How good is Kolibri-1?
All published scores are Aleph Alpha's own: 96.9 on AIME 2025, 84.3 on GPQA-Diamond, 66.4 on SWE-Bench Verified, and leadership claims over Qwen3.6-35B-A3B, Nemotron 3 Super, and Mistral Small 4. No independent benchmark of the model was available at release.
Where was Kolibri-1 trained?
The model card lists 768 NVIDIA B200 GPUs across 96 HGX nodes for pre-training, running 21 days at about 392,000 GPU-hours, with training infrastructure in Germany and Finland, according to Aleph Alpha's disclosures and press coverage.

Sources

  1. Kolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha
  2. Aleph-Alpha/Kolibri-1 model card — Hugging Face
  3. Kolibri technical report (PDF) — Aleph Alpha
  4. Aleph Alpha's Kolibri Is No Match for the Open-Weight Leaders — Trending Topics
  5. Cohere & Aleph Alpha: Transatlantic Sovereign AI — Cohere

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →