Aleph Alpha Releases Kolibri: 78B Mixture-of-Experts Model Optimized for European Languages
Aleph Alpha released Kolibri, an Apache 2.0 open-weight model with 78.1B total parameters and only 3.46B active per token. Featuring a custom UniBPE tokenizer, it cuts token counts for compound languages like German while matching frontier context scales.

Why it matters
You can deploy an open-weight, EU-compliant Mixture-of-Experts model that reduces token consumption and context bloat when processing compound European languages.
TL;DR
- 01Kolibri achieves high inference speeds by evaluating only 3.46B parameters per token from its 78B pool.
- 02The UniBPE tokenizer cuts token usage for German legal and technical text by up to 15%.
- 03Interleaving sliding window and full attention layers enables context lengths exceeding 1M tokens.
Key facts
- Total Parameters
- 78.1 billion
- Active Parameters Per Token
- 3.46 billion (4.4%)
- Native Context Window
- 262,144 tokens (validated up to 1,048,576)
- FP8 Weight Footprint
- Approximately 78 GB
- License
- Apache 2.0
Architecture Breakdown
Kolibri utilizes a sparse Mixture-of-Experts layout designed for on-premise European deployments:
- Parameters: 78.1 billion total, with 3.46 billion active per token (4.4% compute load).
- Structure: 50 layers, 384 experts plus 1 shared expert, routing to top-6 per token.
- License and Weights: Released under Apache 2.0; FP8 weights occupy approximately 78 GB.
Efficient UniBPE Tokenization
Standard English-centric tokenizers fragment compound European nouns into tiny subwords. Aleph Alpha implemented UniBPE—pairing bottom-up BPE merging with Unigram scoring. On legal German text (the German Constitution), Kolibri used 15% fewer tokens than GPT-5's o200k_base tokenizer while tying its token density on standard English benchmarks.
Long-Context and Native Reasoning
To manage context windows up to 1,048,576 tokens without quadratic compute explosions, 40 of its 50 layers use a 512-token sliding window, with full attention activated every fifth layer. On the RULER benchmark at 1M tokens, Kolibri's base model scored 63.2 (outperforming Qwen3.5 35B-A3B at 57.5). Furthermore, its training corpus targeted native non-English reasoning paths, avoiding translation degradation during logic steps.
Try it in 2 minutes
from transformers import AutoModelForCausalLM, AutoTokenizer
# Load Kolibri open weights under Apache 2.0
tokenizer = AutoTokenizer.from_pretrained("Aleph-Alpha/Kolibri-78B")
model = AutoModelForCausalLM.from_pretrained("Aleph-Alpha/Kolibri-78B", device_map="auto")python
✓ When to use
- Enterprise applications requiring strict adherence to European data sovereignty and the EU AI Act.
- High-volume document intelligence across German and English legal or technical texts.
- Workloads requiring long-context reasoning up to 1M tokens with low inference compute costs.
✕ When NOT to use
- Deployments on single 8 GB or 16 GB GPUs without offloading, as 78 GB VRAM is required to host the full model.
- Pure English short-form generation tasks where smaller 7B–14B dense models suffice.
- Environments where maintaining a 128k-token custom tokenizer causes pipeline incompatibilities.
What to do today
- Inspect tokenizer compression ratios using Hugging Face tokenizers with Kolibri weights.
- Evaluate Kolibri in vLLM or Hugging Face Transformers for sovereign European language pipelines.
- Leverage 78 GB FP8 weights to run full 78B capabilities across unified VRAM setups.
What the community says
“People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good.”
“It depends on the use case, but I believe SLMs can do great with smaller tasks. People got so attached to 'general purpose' that they forgot software can be designed for specific, smaller use cases.”