Long context on 6 Sparks: GLM 5.3 full now fits 256K tokens on our engine, and we're working on OS-level fixes that could push it much further. 1M? Cached follow-ups on a 32K conversation come back in ~0.6 s in testing (was ~19 s).

Oct 1, 2026 · 10:37 PM UTC

3
1
12
794
Sort replies: Relevant Recent Liked
Replying to @CK2084
whats the speed
1
92
On 6 Sparks, single request: prefill ~1,800-1,930 tok/s (8K-32K), decode ~35 tok/s prose, ~46 code, ~67 structured. Still early though. Lots of work to do before I ditch vLLM entirely.
1
1
9
250
Replying to @CK2084
Please please stop. I don't need to buy 2 more sparks.
1
1
104
Ha, sorry Jim. The same engine now runs on just 4 Sparks too, and it already beats our published 4-Spark build on decode.
1
21
Replying to @CK2084
Nice work. Glad someone is working on this Is that NVFP4?
1
176
Thanks! It's EXL3 at 3.25 bpw — GLM 5.3 full across 6 DGX Sparks. We tested NVIDIA's official NVFP4 too, and EXL3 came out ahead for us on speed and KV cache.
1
1
63