hi
@ggerganov! just added support for dflash on upstream llama.cpp in atomic chat - thanks a lot for releasing it
DFlash makes Qwen 2.2x faster with no quality loss!
We ran the same Qwen3.6-27B locally three ways on one RTX 6000: baseline, MTP, DFlash. The tasks only differ in one thing - how predictable the next word is: quicksort, describe a file in JSON, a logic puzzle, a sci-fi story.
Outputs:
Baseline: 44 tok/s · 1.00x
MTP: 65 tok/s · 1.45x · 71% accepted
DFlash: 98 tok/s · 2.20x · 30% accepted
Baseline writes one token per step. MTP works inside the model itself and guesses 3 tokens ahead. DFlash is a separate small model that writes 15 tokens at once, and the big model only checks them. In JSON the same words repeat all the time, so most guesses were right: 152 tok/s, 3.4x speedup. In the story 9 guesses out of 10 were wrong. DFlash did all that extra work for nothing and became slower than baseline: 42 vs 44 tok/s. MTP guesses only 3 tokens, so a wrong guess costs very little: 46 tok/s and the win in that round. The output is identical in all three modes - DFlash is the pick for tasks with predictable output, like coding, and for chat and creative writing MTP works better.
DFlash is now natively integrated into Atomic Chat on llama.cpp - speed up your Qwen models!