GLM 5.3 Flash on TensorFold v1.4 is out 🔥
This is a MAJOR release, with a lot of changes, additions, and fixes, starting with the quant.
- NEW EXL3 quant optimized for TensorFold!
- Same size, same speed, better quality.
3x DGX Sparks support:
- New TP=3 path
- Around 6M KV cache
- 77 tok/s on prose, single stream
- 146 tok/s on prose, 4 concurrent stream
- Prefill up to 2064 tp/s
New features / fixed issues:
- Concurrency issues were completely fixed!
- Improvements and fixes in tool calling.
- Stability and diagnostics improvements.
Special thanks to
@YasaarBiladama for letting me use his 2x RTX 6000 PRO server to create the my EXL3 quant!
mia-ai.net/models/GLM-5.3-Fl…