Aleph Alpha released Kolibri on 3 October 2026 as a text-only mixture-of-experts model with a 1M-token context window and four reasoning effort levels. It suits public-sector, industrial and aerospace teams that must self-host under European data rules, not buyers who want the highest open-model score.
Kolibri scores 96.9 on AIME 2025 in Aleph Alpha's own evaluation setup while activating only 3.46B of its 78.1B parameters per token. It is an English-German open-weight model under Apache 2.0, trained to decline answers when the supplied documents do not contain them.
Where it sits
- 66.4%% solvedSWE-bench VerifiedHigher is better#25 / 32peer median 76.8%per source, see benchmark scores
- 84.3%% correctGPQA DiamondHigher is better#31 / 50peer median 88.3%per source, see benchmark scores
Ranks are against GA models on HokAI that publish the same figure; ties share a rank.
Provider: Aleph Alpha · Family: Kolibri
More about Aleph Alpha on HokAI
Context window: 1,048,576 tokens
Input modalities: text, tool-calls · Output: text, tool-calls
About Kolibri
Kolibri is an English-German mixture-of-experts Transformer released by Aleph Alpha on 3 October 2026, the Day of German Reunification, with its weights on Hugging Face under the Apache 2.0 license. It has 78.1B total parameters and 3.46B active per token, drawn from 384 experts of which 6 route each token, across 50 layers that are all mixture-of-experts with one shared expert. Aleph Alpha built it for sovereign work in regulated sectors such as public administration, industrials and aerospace. It follows Kolibri Origin, a 30.6B-parameter internal model that finished pre-training on 11 June 2026 and never had a public release; Kolibri finished pre-training on 11 September.
Every benchmark below is vendor-reported, run in Aleph Alpha's own evaluation setup at the highest reasoning effort, and none has been reproduced by an independent evaluator at the time of writing. Kolibri scores 66.4 on SWE-bench Verified, 85.9 on LiveCodeBench v6, 96.0 on AIME 2026, 21.5 on Humanity's Last Exam, 80.0 on MMLU-Pro, 78.1 on IFBench and 27.7 on Terminal-Bench 2.1. Across the vendor's overall averages it reaches 75.5 in English and 70.8 in German, ahead of GPT-OSS 120B (72.3 and 70.2), Nemotron 3 Super 120B-A12B (73.0 and 67.9) and Qwen3.5 35B-A3B (74.7 and 69.8). It trails Qwen3.8 27B at 80.2 and 79.9, and Qwen3.8 27B is a dense model that spends 27B parameters on every token. On agentic coding the gap is wider: Qwen3.8 27B reaches 76.8 on Terminal-Bench 2.1 and 72.6 on SWE-bench Verified.
The attention design keeps long contexts affordable: of the 50 layers, 10 attend over the full context and 40 use a 512-token sliding window. Training ran in three stages on 768 B200 GPUs: 20T tokens at a 16k sequence length over 21 days, 3.44T tokens of mid-training at 64k, and 200B tokens of long-context adaptation at 256k. The longest trained length is 262,144 tokens, which is also the recommended working limit. Reaching 1,048,576 tokens needs an explicit serving override. The launch post charts LongBench Pro and AA-LCR results but publishes no recall figure at 1M tokens, so treat anything beyond 262,144 as lightly evidenced.
Kolibri takes text and tool definitions in and returns text and tool calls. It has no image, audio or video input. Reasoning effort is selectable per request at four levels (none, low, medium, high), which lets a team trade cost and latency against answer quality. For agent work the vendor reports 94.7 on Tau2-Bench Telecom, 69.9 on Retail and 76.7 on Airline, 61.4 on BFCL v4 overall and 29.4 on BrowseComp. Tool calling and reasoning output are parsed by dedicated vLLM parsers named kolibri1.
HokAI found no hosted per-token price from the vendor for Kolibri, so there is no figure to build cost examples from. The weights carry no license fee under Apache 2.0, which moves the cost to your own GPUs. Enterprise deployment and model specialization are sold through Aleph Alpha's sales team with no public price list.
Deployment is self-hosted. The checkpoint is Aleph-Alpha/Kolibri-1 on Hugging Face, and it needs the aleph-alpha-inference package, which provides the Kolibri plugin for vLLM; the vendor ships a container image and a pip package that also installs the vLLM version it supports. The vendor serves it with an fp8 key-value cache. In the vendor's own sizing test on two H100 GPUs, the 78B design handled 18 concurrent 256k-token requests where a 123B variant handled 3, and it decoded 28% faster. The launch post names no hosted API or cloud marketplace listing, and no minimum VRAM for the released checkpoint.
The safety story is grounding rather than refusal tuning. Kolibri was trained on examples where the right answer is "I don't know", plus Aleph Alpha's Merlin-Arthur procedure, a three-player setup that generates negative examples from the training documents. On the public AA-Omniscience set it abstains instead of answering wrong on 44% of items, against 15% for Kolibri Origin, and its own grounding score is 0.23 where the earlier model scored 0. The same table shows the limit of the approach: its AA-Omniscience Index of -32.8 is well below Qwen3.8 27B at -9.5 because accuracy stays low at 14.8. No system card, red-team partner list or refusal-rate benchmark was published with the launch.
Pick Kolibri if you must run a model on your own hardware under European law, need strong German alongside English, and want an assistant that declines to answer when your documents do not contain the answer. It suits document question answering, agentic retrieval and tool-driven workflows in public sector, aerospace and manufacturing settings, where the vendor reports gains on its internal customer-proxy suites. Pick something else if raw quality per request matters most: Qwen3.8 27B scores higher on the vendor's own table, and Kimi K3 or DeepSeek V4 suit teams that want frontier-scale open weights with hosted access. Avoid it for image, audio or video work and for terminal-heavy coding agents.
Pre-training used 20T tokens: about 62% English, 21.3% German (roughly 4.3T tokens) and 14% code. Aleph Alpha built a 2.4T-token unique German pool, 80% curated or generated in-house from Common Crawl and rephrased organic documents and 20% from open datasets, and used translated text for only 6% of the overall mix. A bilingual UniBPE tokenizer with a 128,000-token vocabulary compresses German at 4.90 bytes per token on FineWeb-2, against 3.72 for DeepSeek V4 and 3.28 for Kimi K3. Post-training used 268B tokens of supervised data and reinforcement learning over more than 1.2 million tasks. The knowledge cutoff is 18 June 2026. Training took place in Germany and Finland, and the vendor says it designed the model with the EU AI Act, the General-Purpose AI Code of Practice and the GDPR in mind. Because the weights run on your infrastructure, data retention is set by the operator.
Against Kolibri Origin, the vendor's internal customer-proxy suites show large gains in three months: German public sector 0.54 to 0.75, aerospace 0.14 to 0.59 and semiconductors 0.35 to 0.80. These suites are internal and unpublished, so they show direction, not a ranking against other labs. Aleph Alpha says it is now considering a larger scale-up on the same pipeline.
Screenshots

Pricing
No hosted per-token API price exists for this release. Downloading the weights costs nothing, so the bill is your own compute. Enterprise deployment help is quoted by sales.
| Tier | Rate |
|---|---|
| Open weights (Hugging Face) | Apache 2.0, self-hosted, no license fee |
| Enterprise deployment and specialization | Contact Aleph Alpha sales; no public price |
Key Features
- Sparse mixture-of-experts: A pool of 384 small experts with 6 routed per token keeps compute near a 3B model while holding the knowledge capacity of a much larger network.
- Mostly local attention: Only 10 of 50 layers attend over the full context; the other 40 use a 512-token window, which bounds memory in long prompts.
- Four reasoning effort levels: Callers choose none, low, medium or high effort per request to trade cost and latency against answer quality.
- Bilingual UniBPE tokenizer: A 128,000-token vocabulary that follows German word structure and reaches 4.90 bytes per token on German web text.
- Merlin-Arthur abstention training: Synthetic negative examples teach the model to say "I don't know" when the documents do not contain the answer.
Pros
- Cheap to serve for its quality because only a small slice of the network runs per token.
- Open weights under a permissive license remove license fees and vendor lock-in for on-premise deployments.
- The abstention behavior is a design goal, not a prompt trick, which matters for regulated document work.
Cons
- Vendor-only benchmarks in a self-run setup, with no independent score yet.
- Trails Qwen3.8 27B on the vendor's own overall averages in both languages.
- No hosted API, no system card and no published hardware requirements at launch.
Benchmarks
- IFBench: 78.1% vendor-reported · 05 Oct 2026 — How precisely the model follows detailed instructions, % passing.
- MMLU-Pro: 80% vendor-reported · 05 Oct 2026 — A harder version of the 57-subject knowledge exam, % correct.
- AIME 2025: 96.9% vendor-reported · 05 Oct 2026 — Competition-level maths problems from the 2025 exam, % solved.
- AIME 2026: 96% vendor-reported · 05 Oct 2026 — Competition-level maths problems from the 2026 exam, % solved.
- GPQA Diamond: 84.3% vendor-reported · 05 Oct 2026 — PhD-level science questions that are hard to search for, % correct.
- LiveCodeBench v6: 85.9% vendor-reported · 05 Oct 2026 — Recently published coding problems, % solved.
- SWE-bench Verified: 66.4% vendor-reported · 05 Oct 2026 — Real GitHub issues fixed end to end, % solved.
- Terminal-Bench 2.1: 27.7% vendor-reported · 05 Oct 2026 — Multi-step tasks completed in a real command line, % solved.
- Humanity's Last Exam: 21.5% vendor-reported · 05 Oct 2026 — Expert-written questions across many fields, % correct.
A benchmark is an exam, not the job. Scores transfer unevenly between tasks, so weigh the one closest to your workload and read every figure with its source.
Frequently Asked Questions
What does Kolibri actually cost?
Aleph Alpha publishes no per-token API price for Kolibri. The weights are free to download under Apache 2.0, so your cost is the GPUs you run them on. Enterprise deployment and specialization are quoted through sales, with no public price list.
What separates Kolibri from Qwen3.8 27B?
Qwen3.8 27B wins on quality: 80.2 versus 75.5 on the vendor's English average and 79.9 versus 70.8 in German. Kolibri wins on serving cost, because it activates 3.46B parameters per token where the Qwen model is dense, and on abstention training, which makes it decline when your documents lack the answer. Pick Qwen for raw quality, Kolibri for cheap on-premise German and English work with grounded answers.
Is Kolibri open source?
Kolibri is open-weight under the Apache 2.0 license, which allows commercial use, modification and redistribution. The checkpoint Aleph-Alpha/Kolibri-1 is on Hugging Face. You also need the vendor's aleph-alpha-inference package, because stock vLLM does not include the Kolibri plugin.
Does Kolibri train on your data?
There is no hosted service to train on your data: you download static weights and run them yourself. The vendor says on-premise operation avoids sending internal data to third-party inference services. No retention policy exists because no hosted API was announced.
Who should skip Kolibri?
Skip it if you need image, audio or video input, because it is text-only. Skip it for terminal-heavy coding agents, where its Terminal-Bench 2.1 score trails Qwen3.8 27B by a wide margin. Teams without GPU capacity to self-host should look at hosted models instead.