Long context on 6 Sparks: GLM 5.3 full now fits 256K tokens on our engine, and we're working on OS-level fixes that could push it much further. 1M? Cached follow-ups on a 32K conversation come back in ~0.6 s in testing (was ~19 s).
3
1
12
794
Nice work. Glad someone is working on this Is that NVFP4?
1
176
Thanks! It's EXL3 at 3.25 bpw — GLM 5.3 full across 6 DGX Sparks. We tested NVIDIA's official NVFP4 too, and EXL3 came out ahead for us on speed and KV cache.
1
1
63
I don't know the state of play now, since the EXL3 folks very rarely share Prefill with decode in Tweets/screenshots.. But EXL3 has/had quite a bit worse prefill than NVFP4. Do you still have the receipts from when you tried NVFP4?
2
45
Replying to @alexellisuk
Here’s what that’s looking like

Oct 2, 2026 · 4:33 PM UTC

2
1
35
Sort replies: Relevant Recent Liked
Replying to @CK2084
Thanks a lot for sharing that. I’d be interested to see a RigMark run on one of your recipes or on your engine. Took around 4-5 mins on GLM 5.3 Flash on TP4. The “full” model is also interesting for us for that cyber capabilities github.com/alexellis/rigmark
71
Replying to @CK2084 @alexellisuk
Just to confirm: is the ~1.9K tok/s prefill from NVFP4, while EXL3 TP6 is around ~1.0K? I may have read the screenshot backwards.
1
45
Other way round: ~1.9K is our custom from-scratch EXL3 engine; ~1.0K is stock EXL3 TP6. NVFP4 was ~1.2K but only had a 219K KV pool, so it wasn't worth it for us.
1
1
35