7 hours ago, I wasn't confident prose would move.
It moved. 1.26x → 1.36x, up 8-10%.
DeepSeek 4.1 Flash on TensorFold is Live! (Still some bugs to resolve)
Same day, G13 build on 2x DGX Spark:
prose 40 → 44.3 tok/s (32.5)
single stream 83 → 85.1 tok/s (kit 32.2)
4 streams 93.5 → 96.6 (kit 37.6)
code 79-81 → 82.8 (kit 42-45)
5 rewrites, 2 were faster on the bench and slower inside the CUDA graph so they went in the bin but stayed in the docs. Fused MoE chain saved 9-46us a layer on its own and cost +0.7ms in the graph.
Quality hasn't moved: 99.6% agreement, MMLU 87.5%.
Still a little more on the table for this
Anyone else had a kernel win in isolation and lose in the graph?