Sort replies: Relevant Recent Liked
Replying to @ggerganov
This is great to see, I think the largest barrier with these different implementations is the requirements for different models and determining what actually yields benefit on your own hardware. For example, trying to run Gemma E4B MTP on a base M3 actually hurts performance.
1
3
1,868
Replying to @ggerganov
DFlash doubled the speed on my AMD Strix Halo.
DFlash doubled the speed of my AMD Ryzen AI Max+ 395. Qwen 3.6 27b: 12.6 t/s --> 24.7 t/s But, I get 80 t/s using Qwen 3.6 35B-A3B with MTP.
283
Replying to @ggerganov
Damn so many options... regarding qwen models any big difference when comparing dflash and MTP? I think dflash is supposed to be faster but at what cost if any?
1
417
Replying to @ggerganov
thank you from all ⚡️ 3090 users!
3
449
Replying to @ggerganov
Looking forward to @lmstudio
159
Replying to @ggerganov
Supercool, Llama.cpp is evolving everyday! I would really love to see some updates on the Android support and inference optimizations there. The Hexagon NPU backend there works, there's Vulkan GPU acceleration too, but seems like there's room for optimization still! GGUF leads tho!
15
Replying to @ggerganov
Another speculative decoding technique landing in llama.cpp is great news for local model speed.
335
Replying to @ggerganov
the number nobody watches is acceptance rate. ran spec decode local on llama.cpp expecting a free 2x. draft was close but not matched to the workload, tokens kept getting rejected, ended up slower than plain decode. whole speedup lives in how often the draft guesses right.
520
Replying to @ggerganov
that is huge for local inference speed, nice work.
447
Replying to @ggerganov
Very cool! Which models have shown the most promise so far with DFlash?
425
Replying to @ggerganov
hi @ggerganov! just added support for dflash on upstream llama.cpp in atomic chat - thanks a lot for releasing it
DFlash makes Qwen 2.2x faster with no quality loss! We ran the same Qwen3.6-27B locally three ways on one RTX 6000: baseline, MTP, DFlash. The tasks only differ in one thing - how predictable the next word is: quicksort, describe a file in JSON, a logic puzzle, a sci-fi story. Outputs: Baseline: 44 tok/s · 1.00x MTP: 65 tok/s · 1.45x · 71% accepted DFlash: 98 tok/s · 2.20x · 30% accepted Baseline writes one token per step. MTP works inside the model itself and guesses 3 tokens ahead. DFlash is a separate small model that writes 15 tokens at once, and the big model only checks them. In JSON the same words repeat all the time, so most guesses were right: 152 tok/s, 3.4x speedup. In the story 9 guesses out of 10 were wrong. DFlash did all that extra work for nothing and became slower than baseline: 42 vs 44 tok/s. MTP guesses only 3 tokens, so a wrong guess costs very little: 46 tok/s and the win in that round. The output is identical in all three modes - DFlash is the pick for tasks with predictable output, like coding, and for chat and creative writing MTP works better. DFlash is now natively integrated into Atomic Chat on llama.cpp - speed up your Qwen models!
75
Replying to @ggerganov
thank you!
361