GLM-5.3 on 6x DGX Spark (TP6): boot time cut from ~250 s to 161 s.- Direct I/O weight loading with a thread pool- Skip redundant expert passes during loadProse decode stayed at 33.8 tok/s.Recipe: github.com/adapt-ai-systems/…

Sep 28, 2026 · 9:15 PM UTC

1
4
154
Sort replies: Relevant Recent Liked
Replying to @CK2084
skipping redundant expert passes during load is the underrated part, MoE checkpoints carry a lot of dead weight if the loader doesnt know which experts actually get routed to at startup. thread pooled direct IO gets you the rest but only after that pruning is done first
3