Open-Source ML Engineer at 🤗 Hugging Face | Creator of oMLX | I build the tools I wish existed for my Mac, then open-source them. junkim.dot@gmail.com

Seoul
Pinned Tweet
I'm happy to announce that I've joined Hugging Face. What started as a personal project back in February is now something I get to work on full time. Local AI has grown explosively this year. I've said this since the early oMLX releases: I want my friend who bought a MacBook yesterday to be able to run AI on it today. I believe MLX has that potential. Apple Silicon is the easiest entry point for a regular person to get started with AI, with no complicated hardware to assemble. And the open source community, including Hugging Face, has been growing that potential. That support has already made a huge difference. Hugging Face is the best place for me to support oMLX and the MLX community with everything I have. For anyone getting into local AI, the first step usually starts at Hugging Face. Mine did too. I'm proud that I get to work at that entry point, where new ideas can be tried out. oMLX stays exactly where it is, under the same Apache 2.0 license in the same repository, and I'll keep leading the project, same as before. What changes is that I can now spend far more time on it, move faster, and build something sustainable for the long term together with all the contributors who have put so much into it. I also want to do more for MLX as a whole. oMLX is built on top of transformers, mlx-lm, mlx-vlm and the rest of the ecosystem, and I'm grateful to the people behind them. Rather than keeping everything inside oMLX, over time I want to push work upstream where it makes sense. And where the community needs something that doesn't exist yet, oMLX is a good place to try it first. Thank you to the 264 contributors who have built oMLX with me, and to everyone who filed issues with detailed logs and reproductions so we could fix things. oMLX would not be what it is without you. oMLX continues in the same place, in the same way. Just much faster. HF's announcement: huggingface.co/blog/omlx
132
102
1,115
94,605
Please stop using MLX 4bit affine quants, focused on something else. They are fast but not good enough, especially in larger contexts. I learned the lesson the hard way using them and even thanks to @ggerganov pointing me in the right direction a long time ago.
43
5
177
17,804
After a long wait, oMLX 0.7.0 is out! This release brings faster Qwen3.8, GLM-5.3-Flash and MiMo V2, faster Lightning MTP and DFlash2 at long contexts, and a completely rebuilt memory guard. github.com/jundot/omlx/relea… * Special thanks to @Spangler3000 and @jerry543, who helped a lot with the speed improvements. Performance vs 0.7.0rc1 (oQ4e, Lightning MTP on) - Qwen3.8-Flash-Next prefill on M5 Max: 1,953 -> 2,844 tok/s (+46%) at 16K context. - Qwen3.8-27B decode on M3 Ultra: 46.9 -> 70.3 tok/s (+50%) at 64K context. - GLM-5.3-Flash prefill on M3 Ultra: 413 -> 516 tok/s (+25%) at 16K context. What's new - A completely rebuilt memory guard. It lets oMLX use as much of your memory as it safely can. - DFlash2 support for GLM-5.3-Flash. - Faster image follow-ups for Qwen3.8-Flash-Next with a vision feature cache. - Fixes for the regressions reported against the RC, including ANE prefill, M5 A8 prefill and clustering. Most of the speedups in this release came from contributors. Thanks to everyone who sent PRs and tested dev&RC versions!
27
37
242
17,341
This is what open source looks like when it's working. Most of these gains came from contributors, not me. Thank you for measuring it properly and making the graph, @viticci!
Seeing some WILD improvements to M5 Ultra performance for MLX after just one week. Measured recent oMLX upstream builds, token prefill has essentially doubled across the board with Flash-Next. Only one week has passed. Will need to update my review soon!
2
3
67
6,896
Huge thanks for this. @Spangler3000! Merging your PRs now. Always grateful for how much energy contributors put into open source. My M5 Ultra is still a month out, and this is making the wait harder!
@jundotkim , I published a bunch of PR's for OMLX for m5U optimization for Mimo 2.6 flash , GLM 5.3 flash and qwen 3.8 Next . As always Lossless by design no tricks or benchmax BLUF : M5 ultra is insanely fast . 🔥 The numbers are pretty wild after a day and a half of lossless optimizing. Qwen3.8-Flash-Next Prefill: 5.6k tok/s @ 8k–16k · 5.5k @ 64k · 4.5–4.8k @ 4k Decode: 109–124 tok/s (MTP) Peak mem: ~109 GB GLM-5.3-Flash Prefill: 2.35k tok/s @ 16k & 64k · 2.32k @ 4k Decode: 63 @ 1k · 59 @ 4k–16k · 56 @ 64k (no spec dec) Peak mem: ~187–191 GB MiMo-V2.6-Flash Prefill: 3.3k tok/s @ 8k Decode: 99–118 @ 8k · 89–99 @ 1k (MTP) Peak mem: ~170–173 GB
3
2
45
7,482
oMLX 0.7.0rc1 is out! This release brings faster Qwen prefill & generation, MiMo V2.6, Ternary Bonsai 2, and partial block caching. DFlash now handles concurrent requests together, and Lightning MTP gets faster batch decoding! github.com/jundot/omlx/relea… Performance on M5 Max, 128 GB (Prefill, oQ4e quant) - Qwen3.8-Flash-Next: 1,522 -> 2,007 tok/s (+32%) at 16K context. (Decode, batch=4, oQ4e quant) - Qwen3.8-27B with DFlash2: 56.9 -> 131.5 tok/s (+131%) - Qwen3.8-27B with Lightning MTP: 88.9 -> 136.9 tok/s (+54%) Full benchmark details are in the release notes. New models and features - MCDMA RDMA support for Mac + CUDA deployments, contributed by @ashxhart. - Partial block caching. No more reprocessing thousands of tokens just because they didn't fill a complete cache block. In one test, next-turn prefill dropped from 1,174 tokens to 37. - Ternary Bonsai 2 text and vision support. - MiMo V2.6 image, video, and audio understanding, plus Lightning MTP and DFlash for compatible checkpoints. - Broader MoE expert offload, including Lightning MTP alongside expert offload for DeepSeek V4.1 and GLM-5.3-Flash. This RC also includes the improvements from the dev releases, including Cluster v2, one-click model settings from community benchmarks, and a customizable dashboard. The GDN prefill kernels are adapted from @ddalcu's excellent mlx-serve! After a short round of testing, I'll publish the stable release and keep moving forward!
17
16
203
10,969
Right after I posted about joining HF on GitHub Discussions, my GitHub account got suspended for an unknown reason. I think it's likely an automated false positive, but I'm still looking into it. I'll share an update as soon as I know more.
12
2
75
5,734
After looking into similar cases, it turns out this happens quite a lot. Not exactly what I wanted on the day of a big announcement, but it is what it is. I've reached out through every channel I can, and hopefully it gets resolved soon.
2
22
1,282
My GitHub account has been restored and the oMLX repo is accessible again. GitHub said it was flagged by their automated abuse detection and cleared after manual review. everything should be back to normal. Thanks everyone for the kind messages and patience today!
15
386
My GitHub account has been restored and the oMLX repo is accessible again. GitHub said it was flagged by their automated abuse detection and cleared after manual review. everything should be back to normal. Thanks everyone for the kind messages and patience today! github.com/jundot/omlx
6
8
184
8,605
Great review from @viticci on the M5 Ultra Mac Studio. Glad oMLX could be part of it. Stay tuned for the next oMLX release!
I reviewed the M5 Ultra Mac Studio with 256 GB of RAM. It’s the dream Mac for local AI agents. And I’ve gone local-only with Hermes. My story, with LOTS of interactive charts: macstories.net/stories/m5-ul…
2
2
75
14,922
oMLX 0.7.0.dev4 adds DeepSeek V4.1 CED prefill with up to 79% faster prompt processing on M3 Ultra (opt-in), multi-request Lightning MTP, and plenty of quality-of-life improvements including one-click model settings. github.com/jundot/omlx/relea… You can now apply model settings with one click using over 450,000 community benchmarks on omlx.ai. Pick a top result for your Mac chip and model to use the same settings. Using a customized model? Copy a one-line recipe from a model with the same architecture and paste it into oMLX. There's also a customizable dashboard, global settings reset, and many bug fixes. See the release notes for the full details. Thank you for contributing code and sharing your benchmarks. Your benchmarks now help other users find and apply settings for their models with one click! * This release upgrades mlx-lm and mlx-vlm, with extensive internal changes. If something that worked before breaks, please open a GitHub issue with logs, and roll back to dev2 for now. ** I test as much as I can before each release, but I can't cover every Mac configuration and model. Issue reports on dev builds are a huge help to me.
8
6
108
5,191