AI NEWS 8 min read

Aleph Alpha Releases Kolibri as an Apache-2.0 English–German MoE Model

Kolibri activates 3.46B of 78.1B parameters per token, publishes weights under Apache 2.0, and makes data control, German efficiency, and deployment cost part of the model specification.

By EgoistAI ·
Aleph Alpha Releases Kolibri as an Apache-2.0 English–German MoE Model

Aleph Alpha has released Kolibri, an English–German mixture-of-experts model with 78.1 billion total parameters and 3.46 billion active parameters per token. The weights are available under Apache 2.0. The company is pitching more than an efficient model: it describes a development and deployment chain designed for organizations that need control over infrastructure, data handling, language performance, and regulatory documentation.

The launch arrives with a model card, public weights, a long technical report, and a public summary of training content. That is a stronger evidence package than a benchmark graphic alone, although the central performance claims still come from the developer and need independent replication.

What happened

Kolibri uses 50 transformer blocks with sliding-window and full attention in a four-to-one pattern. Its sparse feed-forward layers route each token to six of 384 experts plus a shared expert. Aleph Alpha reports about 24 trillion training tokens across pre-training, continued training, and long-context extension, with German accounting for more than 20 percent of the mix.

The model supports a context window of up to one million tokens, while the technical report distinguishes that served limit from the 256K length used in the long-context training stage. The model card recommends shorter contexts for latency- or throughput-sensitive deployment. That qualification is useful: maximum acceptance length and dependable operating range are not the same thing.

Kolibri’s UniBPE tokenizer is designed to improve German compression while retaining competitive English efficiency. The report says German web text requires 11.2 percent fewer tokens than with the GPT-5 tokenizer in its comparison. Tokenization is not just a linguistic detail; fewer tokens can reduce context usage, latency, and serving cost for the same document.

Why it matters

“Sovereign AI” is often used as a political label without an engineering definition. Kolibri makes the term more concrete by connecting it to weight access, deployability on controlled infrastructure, documented data processing, bilingual performance, and a pipeline developed and run in Europe.

For a public agency or regulated manufacturer, model quality is only one procurement axis. The organization may need to know where prompts travel, which administrator can access logs, whether weights can run inside a specific boundary, how an update is versioned, and what license permits modification. An Apache-2.0 release makes more of those choices available to the adopter than a closed API does.

The architecture also illustrates a recurring tradeoff in sparse models. Only a small fraction of weights compute each token, but all weights still consume storage and usually memory. Kolibri may offer attractive decode economics on suitable hardware without becoming a tiny model in operational terms.

Evidence

The technical report provides architecture tables, data-stage descriptions, ablations, benchmark groups, throughput methodology, and tokenizer comparisons. Aleph Alpha says Kolibri sits on a quality-versus-serving-cost Pareto frontier among the open models it tested in English and German, including comparisons with Qwen3.6 35B-A3B, Mistral Small 4, and Nemotron 3 Super.

Those results are useful but bounded. The serving comparison uses the developer’s selected hardware, software, sequence lengths, and model set. Composite benchmark averages can hide task-specific weaknesses. A German public-sector evaluation does not establish performance on a company’s own abbreviations, scanned forms, retrieval errors, or tool permissions.

The Hugging Face model card adds deployment guidance and responsible-use constraints. Community interest on GeekNews is an attention signal, not a quality test. At publication time, broad independent production studies were not yet available.

Practical takeaway

Teams evaluating Kolibri should separate five questions: license, data governance, task quality, serving cost, and operational control. Apache 2.0 answers part of the first question, not the other four.

Run a bilingual held-out test with the exact document types and response formats in scope. Report accuracy separately by language and task. Include abstention quality when the retrieved context does not support an answer. Measure token counts as well as request counts, because tokenizer efficiency can materially change a long-document budget.

For deployment, benchmark the intended quantization, batch size, context distribution, and hardware. Record the full memory footprint, not only active parameters. Preserve model, tokenizer, prompt, and serving-version identifiers so an audit can reconstruct an output.

Limitations

Kolibri’s headline benchmark and efficiency results are primarily first-party. One-million-token support should not be interpreted as uniform recall or reasoning across one million tokens. Open weights do not disclose every training example, eliminate licensing review, or guarantee compliance with a regulation. “Sovereign” depends on the complete system—including hosting, administrators, retrieval sources, logs, and update control—not the model alone.

Share this article

> Want more like this?

Get the best AI insights delivered weekly.

By subscribing, you agree to our Privacy Policy. You can unsubscribe at any time.

> Related Articles

Tags

KolibriAleph Alphaopen weightssovereign AImixture of experts

> Stay in the loop

Weekly AI tools & insights.