GLM-5.3 on 6x DGX Spark (TP6): boot time cut from ~250 s to 161 s.- Direct I/O weight loading with a thread pool- Skip redundant expert passes during loadProse decode stayed at 33.8 tok/s.Recipe: github.com/adapt-ai-systems/…
Sep 28, 2026 · 9:15 PM UTC
1
4
154

