germany just dropped a sovereign open weight model kolibri by @Aleph__Alpha runs 3.5b of its 78b parameters per word and its math is kinda ridiculous: 96.9% on aime beats every mixture-of-experts model they tested, even 3x bigger ones. only a dense model doing 8x the work wins anyone can run it on their own servers, it thinks in german, and in their evals it tops every open model its size in english and german models read text in chunks called tokens, and i ran kolibri's chunker (its tokenizer) on the german constitution: it needed 15% fewer tokens than gpt-5's for the same text. "bundesverfassungsgericht" is 6 tokens for gpt-5 and 2 for kolibri. fewer tokens means cheaper, faster german, and more of it fits in what the model can read at once how it works, simply: 1. every layer has 384 tiny specialists, and a router sends each word to 6 of them. so it thinks like a 3.5b model, but it needs the memory of a 78b one: about 78 gb, which means 2 big nvidia gpus (h100s) or 1 h200 2. most layers only look at the last 512 tokens, and every 5th layer looks at everything. that's how it can read 1 million tokens (a few thick books) without it costing a fortune 3. it reasons in german. their team found that a little german reasoning data is worse than none: the model's german thoughts go in circles and never finish. so they made about 800k german reasoning examples and gave it a lot 4. it's trained to say "i don't know". they play a game with it where parts of the documents are hidden, sometimes to help it and sometimes to hide the evidence, and it has to tell which. when it didn't know an answer it admitted it 44% of the time. qwen3.5 did 11% where it's weaker: answering from memory, using tools over a long back and forth, and coding agents, where qwen models are ahead. and to run it you need aleph alpha's add-on for vllm, a popular open source server for running models if you have german documents and need to keep them on your own hardware, this is a big deal. huge congrats to everyone at aleph alpha, my good friend @MichaelLHofmann included!! i wrote up how it works, the benchmarks, how to run it and when to use it: tej.as/blog/aleph-alpha-koli…

Oct 3, 2026 · 9:08 AM UTC

16
20
173
8,749
Sort replies: Relevant Recent Liked
worth splitting the two numbers: 3.5b per word is the compute bill, 78b total is the memory bill. sparsity buys speed, not a smaller deployment. size the box for 78 and enjoy the price of 3.5.
1
1
284
yes, this is the trap with moe headlines. all 78b have to sit in memory even though only 3.5b work on each token, so it's a data-center gpu or two, never a laptop. cheap per token once it's up though
236
Great breakdown. The number I'm proudest of isn't in the math column though: Kolibri says "I don't know" when the context doesn't support an answer.
2
8
314
Awesome analysis @TejasKumar_ , hyped to get this into some production apps and your feedback soon!
1
3
209
I already have 2 apps I’m adding it to as we speak. Thank you for your good work and for making my Saturday so much more fun!!
1
3
203
Is the 78 GB at 8 bit? At bf16 it'd be about twice that, and 78 GB on its own would nearly fit one H100.
1
184
yes, 78 gb is fp8, so bf16 would be about 156. and one 80 gb h100 doesn't quite cut it even at fp8, since the kv cache and activations need room on top of the weights. that's why the model card's minimum is 2 h100s or 1 h200
1
169
What a great write-up; many thanks!
1
1
41
🫡
1
31
dense kolibri beating bigger moe on aime is a reminder that active params aren't the whole story
1
1
157
exactly
148
Would be interesting to see if its good at telling jokes - the latter aside :D good to see anything coming out of Europe for a change.
1
1
101
haven't tried jokes yet lol, it needs a data-center gpu and nobody hosts it so far. will report back. and same, more of this from europe please
1
1
94
Well I'm all for European/German sovereignty but isn't this company 90% Canadian
23
Was this the Big Pickle free model in Opencode? I've been using it a lot lately after running out of Claude and Codex tokens
35
@grok How does this model compare to the current frontier models?
1
62
We’ve got it up and running as well for anyone to experience it. No GPU. No setup. Just try it. tesseracted.com/kolibri-1-ch…
Dear @Aleph__Alpha team - thank you for making Kolibri-1 open. We care deeply about sovereign AI, and launching it on German Unity Day makes today feel especially fitting. As a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days. No GPU. No setup. Just try it. tesseracted.com/kolibri-1-ch… Open Models move us all forward. We’re rooting for you. 🇩🇪 Tesseracted Labs GmbH. tesseracted.com #alephalpha #openweightmodels
1
3
102
*thinking in german* > profit ;)
14
Needs a strata quant modification
42