The engine from Sunday's demo is out today, with the code and the weights. And a paper. The paper is about one decision the engine makes every round. A drafter guesses a tree of up to 31 next tokens and the big model checks the whole tree in one pass, so somebody has to decide how big that tree should be. It measures what a check costs on the Spark and cuts the tree to the size that pays best. 49.9 tok/s on the benchmark row. A fixed draft of 8 does 47.1, a fixed 16 does 44.8, and the same weights with no drafter do 13.7. 6 min of video: the cut, the dashboard, opencode doing a task on it, and the caches. Links are in the last post.
3
6
38
7,987
Checking 16 guesses costs the same as checking 8. At 1k context, one verify pass on the Spark takes 75.5 ms for 8 rows, 74.4 for 16, 83.1 for 24 and 87.8 for 32. The cost goes up in steps, with a flat step between 8 and 16. My old router used one flat price for the verify, and the price was wrong. Most requests sat on the wide draft and paid about 6.8 ms extra per step. StairCut measures the staircase per context length, from 1k to 32k. At 32k the extra rows cost more. Then every round it builds a candidate tree and cuts it to the node count with the most expected committed tokens per millisecond, at the measured price, snapping to 7, 15, 23 or 31 nodes. It also decides per round whether the lookup drafter, the one that copies from the context, is enough on its own. Same 25 texts forced through every setting: 90.5 tok/s average against 82.0 for a fixed block of 16 and 59.3 for a fixed 8. Above block 16 on 23 of the 25 texts, above block 8 on 22.
2
1
9
388
The dashboard can show precisely what the current request is doing atm, its decode and prefill speeds
1
1
7
230
Coding agents resend the whole conversation on every turn. At 128k that resend was the part you wait for, and now you wait for it once. Cold first read of a 128k prompt: 208.3 s. Next turn of the same conversation, 1.1 s. 64k goes from 93.7 s to 0.6 s, and 32k from 44.4 s to 0.3 s.
1
1
8
207
One line installs it on a Spark: curl -fsSL raw.githubusercontent.com/0x… | bash It checks the GPU, downloads the base model, the NVFP4 weights and both drafters, then starts an OpenAI-compatible server with a dashboard. I only have the one Spark. Other CUDA GPUs are best effort and I have not run it on any of them.
1
1
8
289
Sort replies: Relevant Recent Liked
Replying to @0xBakeer
Awesome work you are doing man!
1
1
61
Thank you Niko 🙏
30