Wouldn't you like to know

Chuck 208 retweeted
Unfortunately I think I can see how you get to the new $5,000 & $7,000 MSRP pricing on the Sparks Priced pulled from what I could get at local Microcenter, Newegg, Amazon or Ebay
3
2
12
1,228
Chuck 208 retweeted
Welcome to Kindle Spark OS! TL;DR: Add 4GB of RAM to any DGX Spark & faster! @CK2084, @coffeedev and I figured out how to get 5-10% extra decode and 4GB RAM out of a Spark. We had to replace NVIDIA's OS to do it, but we made it safe and easy. github.com/kindlingai/kindli…
36
48
408
59,711
Running two kernel efforts side by side with @mmastrac: our agents talk to each other over an encrypted channel and trade measurements, not code. Today they found our NVIDIA Sparks read memory ~5% faster than his ASUS GX10s. Fun times.
1
1
8
531
Spun up a second lane today: the same custom GLM 5.3 full engine on just 4 DGX Sparks. First tests: GPU-to-GPU comms 11-21% faster per step than our 6-node setup. More people own 4 Sparks than 6, so this one matters.
2
5
509
Chuck 208 retweeted
GLM 5.3 Flash on 2x sparks is about to get even faster with kindling-spark-os. We had to dig down into the OS to unlock the last bits of performance. It's a new, reversible boot that gives supported recipes a nice 5-10% boost and more RAM!
3
5
56
2,902
Long context on 6 Sparks: GLM 5.3 full now fits 256K tokens on our engine, and we're working on OS-level fixes that could push it much further. 1M? Cached follow-ups on a 32K conversation come back in ~0.6 s in testing (was ~19 s).
3
1
12
788
GLM 5.3 FULL on 6x DGX Spark, update: our from-scratch engine now prefills ~1,930 tok/s at both 8K and 32K. Last night it was ~1,100 at 8K and ~820 at 32K. Same 3.25 bpw weights, quality gate passed on every step. Decode is next.
3
2
20
905
Our GLM 5.3 (full) build for 6× DGX Spark is up in the kindling org, working with @mmastrac to get the best Spark recipes for this model. Decode: ~32 tok/s prose, ~46 code, ~60 structured Prefill: ~1,100 tok/s KV cache: 767K tokens github.com/kindlingai/glm-5.…
6
4
25
1,115
Chuck 208 retweeted
Add another fact to DGX Spark lore.
PSA if you’re running DeepSeek/Qwen Flash with SSD-backed Engram/PLE tables on a DGX Spark: check your NVMe interrupt coalescing. Found NVIDIA ships a boot service that enables it specifically for Samsung, Kioxia and Micron drives. My Samsung PM9E1s default to OFF. NVIDIA turns it ON. Tested both Samsungs. 20,000 random 4K reads, QD1, direct I/O: coalescing ON: ~200µs OFF: 55–57µs back ON: ~200µs again That’s ~3.5× the latency because of a boot-time setting. If your token generation is waiting on uncached Engram lookups, this can fuck up your latency. A PCIe 5.0 sticker and a giant sequential bandwidth number won’t save you. My first 4 sparks have a no name PCIe 4.0 SSD and my newest 2 came with a Samsung PCIe 5.0. But when I was actually testing SparkNest out, it was preferring the old devices for engram paging into the mmapped backend… because the Samsung drives get configured to buffer 8 read response interrupts… stupid… @NVIDIAAI time to delete that customization script now the new models need fast latency random read. Check: sudo nvme get-feature /dev/nvme0 -f 8 -H The service is nvidia-nvme-interrupt-coalescing.service. This is measured storage latency, not a claim of 3.5× more tokens/sec. Cache hits, batching and prefetching matter. But check this before blaming the model or buying another SSD. For low free memory hosts the impact is higher cuz the mmapped PLE can’t hold the whole engram table.
1
3
34
8,436
Chuck 208 retweeted
A lot of great work by the community @O80925253 @majewskizby @mmastrac on GLM 5.3 Flash. Please show these engineers some appreciation. Please follow and get them a coffee if you can. These guys broke a huge barrier. There are many more engineers to highlight, and don’t be afraid to mention them.
2
2
23
1,083
Chuck 208 retweeted
For anyone building GLM 5.3 Flash on Nvidia's awesome quant (IMO) on consumer hw, I pushed up a sidecar with the original GLM 5.3 Flash weights for the MLA + shared experts and dense MLPs of layers 0–2. If you aren't running datacenter hw, worth it: huggingface.co/mmastrac/GLM-…
3
3
21
1,893
Speed update on our GLM flash Spark build. I gained some ground by removing my clock caps. I probably don’t need them now that I figured out cooling. github.com/kindlingai/glm-5.…
1
6
242
I usually run my dgx sparks clock caps around 1950. On larger models like GLM 5.3 it doesn’t seem to slow things down… but on smaller models like the flash variants it makes a 5-6% difference.
1
91
GLM 5.3 Flash on 6× DGX Spark (TP=6), built on @mmastrac's recipe: • code 120 / prose 66.5 / structured 176 tok/s (RigMark)
• 282 tok/s at 8 code streams
• 887 tok/s steady at 64 streams
• ~5k tok/s prefill at 32k
2
8
984
Phase-aware speculative decoding: GLM's own tokens show whether it's reasoning, writing, coding or calling a tool. We draft 4 tokens ahead on predictable tool calls, fewer on prose, and switch at each boundary. +6.7% on real prompts, +12% on prose. github.com/adapt-ai-systems/…
5
132
Got a nice speed boost by giving the model serving infrastructure the ability to adjust prediction based on token type.
59
GLM-5.3 on 6x DGX Spark (TP6): boot time cut from ~250 s to 161 s.- Direct I/O weight loading with a thread pool- Skip redundant expert passes during loadProse decode stayed at 33.8 tok/s.Recipe: github.com/adapt-ai-systems/…
1
4
154
Chuck 208 retweeted
It's time to let the public test this one out. The fastest TP=4 GLM 5.3 Flash recipe is now ready for wider testing. Thx @CK2084 for help & @ayayalar for testing This one pushes your ConnectX setup to the limit Want to try it out early? Repo: github.com/mmastrac/glm-5.3-…
4
2
32
4,029
GLM-5.3 on 6x DGX Spark (TP6), long context:- Prefill: 1,000 tok/s at 8K- 921 tok/s at 64K- 854 tok/s at 128K- KV pool: 803,968 tokensOnly 15% slower at 128K than at 8K.Recipe: github.com/adapt-ai-systems/…
1
2
161