Long context on 6 Sparks: GLM 5.3 full now fits 256K tokens on our engine, and we're working on OS-level fixes that could push it much further. 1M? Cached follow-ups on a 32K conversation come back in ~0.6 s in testing (was ~19 s).
On 6 Sparks, single request: prefill ~1,800-1,930 tok/s (8K-32K), decode ~35 tok/s prose, ~46 code, ~67 structured. Still early though. Lots of work to do before I ditch vLLM entirely.
Thanks! It's EXL3 at 3.25 bpw — GLM 5.3 full across 6 DGX Sparks. We tested NVIDIA's official NVFP4 too, and EXL3 came out ahead for us on speed and KV cache.