Mark my words, the mlx community is awakening. What holds back local Super Intelligence on Apple silicon isn’t the hardware, it’s the software.
MLX-Serve 26.10.1 is out: speed across the board. Exactly faster. Qwen3.8 27B with its drafter +66% on an M5 Ultra, +28% on an M4 Max, +37% on an M1 Pro. Qwen Flash Next: +11% decode on M4 Max, +51% prefill on M5 Ultra. (up to ~5400 PP!!!!) Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models). New: * GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental) * 🍣 Sushi-format Flash Next packs run directly. * Qwen-Image 2.1 edits pictures from instructions. * Drafting is on by default for every model that supports it. Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle. Special thanks to @ashxhart and @Spangler3000 for pushing ! Benchmarked and tested across 5 machines: htmlhost.jax.workers.dev/ren… GH Release Download + Changelog: github.com/ddalcu/mlx-serve/… Website: mlxserve.com
16
Cowboy Coder retweeted
MLX-Serve 26.10.1 is out: speed across the board. Exactly faster. Qwen3.8 27B with its drafter +66% on an M5 Ultra, +28% on an M4 Max, +37% on an M1 Pro. Qwen Flash Next: +11% decode on M4 Max, +51% prefill on M5 Ultra. (up to ~5400 PP!!!!) Gemma 4, Qwen, LFM2.5, Spark, Bonsai 2 all faster, output byte-identical to 26.9.6 (across 18 models). New: * GGUF models run on our own MLX engine, written in Zig and Metal, instead of the llama.cpp fallback. (experimental) * 🍣 Sushi-format Flash Next packs run directly. * Qwen-Image 2.1 edits pictures from instructions. * Drafting is on by default for every model that supports it. Huge thanks to everyone who filed, tested and fixed this one: @STRML_, @sbusso, @wutang_superfan , lborloz, LXD-8 @alinselea zeeshanhaque21 @Lojza3D, h9q2cyxvgm-ui @kennethrdegraff, @CowboyCoderHQ , codysk, @Beamsters1, @CerebralCoding_, @AjAbsaki , and more.. THANKS ! Please comment if I did not include your name, for visibility, I know people mostly by their GH handle. Special thanks to @ashxhart and @Spangler3000 for pushing ! Benchmarked and tested across 5 machines: htmlhost.jax.workers.dev/ren… GH Release Download + Changelog: github.com/ddalcu/mlx-serve/… Website: mlxserve.com
47
42
388
29,201
Local ai is the way to go, the main downside in my eyes is that you pay every second you aren’t running max tok/sec, because your hardware is an asset, not just your ticket to intelligence, freedom, and privacy. Who else feels my pain?
1
9