MLX-Serve 26.10.1 is out: speed across the board.
Exactly faster. Qwen3.8 27B with its drafter
+66% on an M5 Ultra,
+28% on an M4 Max,
+37% on an M1 Pro.
Qwen Flash Next:
+11% decode on M4 Max,
+51% prefill on M5 Ultra. (up to ~5400 PP!!!!)
Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models).
New:
* GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental)
* 🍣 Sushi-format Flash Next packs run directly.
* Qwen-Image 2.1 edits pictures from instructions.
* Drafting is on by default for every model that supports it.
Huge thanks to everyone who filed, tested and fixed this one:
@STRML_,
@sbusso,
@wutang_superfan , lborloz, LXD-8
@alinselea zeeshanhaque21
@Lojza3D, h9q2cyxvgm-ui
@kennethrdegraff,
@CowboyCoderHQ , codysk,
@Beamsters1,
@CerebralCoding_,
@AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle.
Special thanks to
@ashxhart and
@Spangler3000 for pushing !
Benchmarked and tested across 5 machines:
htmlhost.jax.workers.dev/ren…
GH Release Download + Changelog:
github.com/ddalcu/mlx-serve/…
Website:
mlxserve.com