Run GLM 5.3 Flash EXL3 with TensorFold ⚡️
This is a completely new recipe that ushers a whole new level of performance for
@NVIDIAAI 2x DGX Sparks.
Conservative default for stability:
- 1M context by default
- 2.7M KV cache pool (!)
- Yes, it's not a typo - 2.7M KV in just two Sparks
- 4 concurrent streams by default
Performance:
- 60 tok/s on prose, single stream.
- 108 tok/s on prose, 4 concurrent streams.
~1,950 prefill tok/s for most context lengths.
Stress-tested to handle a variety of workflows.
This is by far the BEST model to run if you have two DGX Sparks. Extremely smooth experience!
Expect further improvements!
Thanks to
@ashxhart for developing such a powerful engine! TensorFold will be used in many of my upcoming recipes.
Get it here:
github.com/MiaAI-Lab/GLM-5.3…