PSA if you’re running DeepSeek/Qwen Flash with SSD-backed Engram/PLE tables on a DGX Spark: check your NVMe interrupt coalescing.
Found NVIDIA ships a boot service that enables it specifically for Samsung, Kioxia and Micron drives. My Samsung PM9E1s default to OFF. NVIDIA turns it ON.
Tested both Samsungs. 20,000 random 4K reads, QD1, direct I/O:
coalescing ON: ~200µs
OFF: 55–57µs
back ON: ~200µs again
That’s ~3.5× the latency because of a boot-time setting.
If your token generation is waiting on uncached Engram lookups, this can fuck up your latency. A PCIe 5.0 sticker and a giant sequential bandwidth number won’t save you.
My first 4 sparks have a no name PCIe 4.0 SSD and my newest 2 came with a Samsung PCIe 5.0. But when I was actually testing SparkNest out, it was preferring the old devices for engram paging into the mmapped backend… because the Samsung drives get configured to buffer 8 read response interrupts… stupid…
@NVIDIAAI time to delete that customization script now the new models need fast latency random read.
Check:
sudo nvme get-feature /dev/nvme0 -f 8 -H
The service is nvidia-nvme-interrupt-coalescing.service.
This is measured storage latency, not a claim of 3.5× more tokens/sec. Cache hits, batching and prefetching matter. But check this before blaming the model or buying another SSD. For low free memory hosts the impact is higher cuz the mmapped PLE can’t hold the whole engram table.