Long context on 6 Sparks: GLM 5.3 full now fits 256K tokens on our engine, and we're working on OS-level fixes that could push it much further. 1M? Cached follow-ups on a 32K conversation come back in ~0.6 s in testing (was ~19 s).
Thanks! It's EXL3 at 3.25 bpw — GLM 5.3 full across 6 DGX Sparks. We tested NVIDIA's official NVFP4 too, and EXL3 came out ahead for us on speed and KV cache.
I don't know the state of play now, since the EXL3 folks very rarely share Prefill with decode in Tweets/screenshots..
But EXL3 has/had quite a bit worse prefill than NVFP4.
Do you still have the receipts from when you tried NVFP4?
Thanks a lot for sharing that.
I’d be interested to see a RigMark run on one of your recipes or on your engine.
Took around 4-5 mins on GLM 5.3 Flash on TP4.
The “full” model is also interesting for us for that cyber capabilities
github.com/alexellis/rigmark
Other way round: ~1.9K is our custom from-scratch EXL3 engine; ~1.0K is stock EXL3 TP6. NVFP4 was ~1.2K but only had a 219K KV pool, so it wasn't worth it for us.