Aleph Alpha Releases Kolibri 1: Sovereign German MoE LLM
Aleph Alpha launched Kolibri 1 on October 3, 2026, introducing an open-weight large language model engineered specifically for German and English applications. Licensed under Apache 2.0, the model represents a strategic push toward sovereign European AI, developed entirely on domestic infrastructure in Germany and Finland and designed for full compliance with the EU AI Act. Kolibri utilizes a mixture-of-experts architecture containing 78.1 billion total parameters, activating approximately 3.46 billion per token. Trained from scratch on 24 trillion tokens, with over one-fifth dedicated to German corpora, the model leverages 768 NVIDIA B200 GPUs to achieve a balance between computational efficiency and multilingual proficiency. The architecture supports a native context window of 262,144 tokens, with validated performance extending to one million tokens through a hybrid attention mechanism that combines sliding-window processing with periodic full-context reviews. Key technical differentiators center on language optimization and regulatory alignment. Kolibri employs a custom UniBPE tokenizer with a 128,000-vocabulary size, specifically calibrated to handle German compound structures, reducing token fragmentation by up to 15 percent compared to leading commercial alternatives. To mitigate hallucination risks in high-stakes environments, Aleph Alpha integrated the Merlin-Arthur protocol, training the model to reliably abstain when supporting evidence is insufficient. Benchmarks indicate the model correctly identifies knowledge gaps 44 percent of the time under rigorous testing, significantly outperforming comparable open-weight systems. Additionally, Kolibri features a German-specific reasoning pipeline, enabling coherent step-by-step logic in native prompts, and includes a configurable reasoning-effort parameter allowing users to scale computational intensity per request. Independent evaluations position Kolibri as a leader among models with similar active parameter counts. It achieves top-tier scores in overall German and English benchmarks, excelling particularly in mathematical reasoning and long-document retrieval-augmented generation. However, the model exhibits trade-offs inherent to its design. Closed-book knowledge retrieval lags behind competitors, multi-turn tool calling and coding benchmarks require improvement, and mid-range context windows show reduced accuracy compared to specialized dense architectures. Deployment requires approximately 78 gigabytes of VRAM, necessitating data-center grade hardware such as dual NVIDIA A100 or H100 units, or a single H200 or B200. The model currently relies on a proprietary vLLM plugin and is not yet available through major hosted inference providers. Aleph Alpha has positioned Kolibri for regulated sectors including public administration, finance, and industrial manufacturing, where data residency, language precision, and auditability are paramount. While unsuitable for resource-constrained environments or multilingual general-purpose tasks, the model establishes a viable framework for sovereign, enterprise-grade AI deployment within the European market.
