Run and train models locally with the Unsloth Desktop app. 🦥 github.com/unslothai/unsloth

San Francisco, CA
Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere Unsloth Desktop is now available on unsloth.ai and GitHub. GitHub: github.com/unslothai/unsloth Blog and Guide: unsloth.ai/docs/desktop
283
647
4,828
869,424
Unsloth AI retweeted
We're co-hosting an open source party with @HuggingFace! 🤗🦥 Come join us, we'll be handing out @UnslothAI merch, showcasing our upcoming features and more! It'll be a super fun night with lots of demos, DJs, food and more. Join: luma.com/OpenTogether Use code: NOSLOTHSHERE to get in.
6
11
113
16,429
You can now run Laya Decision models locally on just 4GB RAM! 🔥 Works on CPU, Mac, Windows, Linux and GPU setups. Serve Laya through a Jev-compatible API via Unsloth Desktop. GitHub: github.com/unslothai/unsloth Guide: unsloth.ai/docs/models/decis…
70
314
2,431
153,102
We got @UnslothAI a DGX Station! @DanielHanChen and @NaderLikeLadder checked out Unsloth’s new @Dell Pro Max with GB300 and talked about what comes next: support for more models, faster quantization, and more efficient reinforcement learning.
30
24
296
24,965
Thanks NVIDIA for the DGX Station! We'll be using it for quantization, testing, training amongst many other things! 🥰
1
31
804
Unsloth AI retweeted
We got @UnslothAI a DGX Station! @DanielHanChen and @NaderLikeLadder checked out Unsloth’s new @Dell Pro Max with GB300 and talked about what comes next: support for more models, faster quantization, and more efficient reinforcement learning.
30
24
296
24,965
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗 Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever. Thanks for all your support!
58
67
1,070
50,748
Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! 🖼️ The 7B model performs on par with Nano Banana 2.0. For higher quality, you can also run Dynamic FP8 on just 6GB of VRAM via offloading. GGUF: huggingface.co/unsloth/Qwen-… Guide: unsloth.ai/docs/models/qwen-…
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: qwen.ai/blog?id=qwen-image-2… - GitHub: github.com/QwenLM/Qwen-Image… - Model Scope: modelscope.cn/models/Qwen/Qw… - Hugging Face: huggingface.co/Qwen/Qwen-Ima…
75
272
2,912
356,051
Qwen-Image-2.1 FP8 and GGUF quants should now run properly in Unsloth Desktop! 💜 Image gen and editing are both supported. GitHub: github.com/unslothai/unsloth
5
7
60
13,269
You can now train and run 500+ models locally with our Unsloth Docker image! 🐳 Use our new GUI or notebooks workflow. No setup required. Works on NVIDIA and AMD. Guide: unsloth.ai/docs/get-started/… GitHub: github.com/unslothai/unsloth
Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere Unsloth Desktop is now available on unsloth.ai and GitHub. GitHub: github.com/unslothai/unsloth Blog and Guide: unsloth.ai/docs/desktop
23
99
706
54,749
Unsloth AI retweeted
Friendly reminder that you can fine tune 500+ open source models in a free Google Colab You can even upload PDFs/CSVs and turn them into usable synthetic datasets. 1. Open the Google Colab below 2. Run the blocks to install Unsloth Studio 3. Choose a model 4. Upload a dataset 5. You're good to go! And you can of course export your model afterwards.
7
74
595
43,608
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
1,099
3,087
28,610
6,922,932
Congrats DeepSeek on another epic release! Hopefully you guys will release smaller models for people to run locally. 🙏🐋 It's great that DeepSeek-V4.1 has 196B engram making it more accessible.
21
21
1,049
80,160
Qwen3.8-27B Unsloth GGUF is now the #1 most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: huggingface.co/unsloth/Qwen3… Guide: unsloth.ai/docs/models/qwen3…
80
124
1,748
93,283
ICYMI, we're celebrating 1 BILLION+ downloads for @googlegemma 💎 How are developers actually using open models? @GoogleDeepMind’s @DynamicWebPaige caught up with devs and collaborators like @UnslothAI and @Qualcomm to hear how they’re building on-device tools, running local fine-tuning, and pushing multimodal breakthroughs.
11
20
189
56,104
Congratulations Google Gemma team on 1 billion downloads! 🤯🔥
17
1,381
Unsloth AI retweeted
ICYMI, we're celebrating 1 BILLION+ downloads for @googlegemma 💎 How are developers actually using open models? @GoogleDeepMind’s @DynamicWebPaige caught up with devs and collaborators like @UnslothAI and @Qualcomm to hear how they’re building on-device tools, running local fine-tuning, and pushing multimodal breakthroughs.
11
20
189
56,104
Unsloth AI retweeted
Local GGUF inference is now up to 3.3× faster at long context lengths
We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction. Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp. Guide: unsloth.ai/docs/models/glm-5… GGUF: huggingface.co/unsloth/GLM-5…
16
4
273
27,088
We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction. Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp. Guide: unsloth.ai/docs/models/glm-5… GGUF: huggingface.co/unsloth/GLM-5…
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
72
140
1,480
233,183
You can now run Unsloth GGUFs locally in one-click via Hermes! ✨ Qwen3.8-27B, Qwen3.8-Flash, DeepSeek-V4-Flash and more are all supported.
Hermes Desktop now sets up local models in one click. It automatically reads your hardware, picks the best model for you, then downloads it and configures the runtime.
28
75
890
73,807
Hermes Desktop now sets up local models in one click. It automatically reads your hardware, picks the best model for you, then downloads it and configures the runtime.
261
322
4,480
415,291
This is super exciting! Go open-source! 😍
1
1
77
3,829
Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗 blogs.nvidia.com/blog/nvidia…
1,715
3,319
27,353
6,124,933
Congrats guys! This is super exciting for open-source and local AI! 💚🤗
58
3,303
Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: huggingface.co/unsloth/Qwen3… Guide: unsloth.ai/docs/models/qwen3…
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: unsloth.ai/docs/models/qwen3… GGUF: huggingface.co/unsloth/Qwen3…
65
128
1,189
115,184
⚡ Fine-tune and quantize AI models directly on NVIDIA Jetson. Our new Jetson AI Lab tutorial shows how to use Unsloth and memory-efficient QLoRA to customize models, export quantized GGUF files, and run them locally with llama.cpp. Follow hands-on examples for: 🔹 NVIDIA Nemotron 3.5 Lightning on Jetson AGX Thor 🔹 Qwen3.5-4B on Jetson Orin Nano Start optimizing: nvda.ws/45YagAq
13
53
381
24,019
Super cool thanks for sharing guys! 💚🦥
1
1
7
1,252