🚀 1,000+ TOKENS/S ON A 1T MODEL! 🚀
We are thrilled to release Xiaomi MiMo-V2.5-Pro-UltraSpeed in collaboration with @TileRT_AI , breaking the 1,000 tokens/s output speed on a 1 Trillion parameter model for the FIRST TIME!
Not wafer-scale integration like Cerebras. Not pure on-chip SRAM chips like Groq. We achieve 1,000 tps on a 1T MoE model using just a SINGLE, STANDARD 8-GPGPU NODE.
Read the full technical deep dive:mimo.xiaomi.com/blog/mimo-ti…
Want to experience the future of real-time AI?
👉 Apply for UltraSpeed now: platform.xiaomimimo.com/ultr…
⏳ Limited-Time Access: Application-based · Jun 8 – Jun 23 (PDT)
💬 Chat Experience: Completely FREE for a limited time — try the blazing-fast web chat now.
⚡ UltraSpeed API: Just 3x the price for a ~10x boost in output experience.
🤝 Enterprise & Large-Scale Needs: business-mimo@xiaomi.com
158
293
2,372
421,162
🔓 And the best part — we're open-sourcing it.
1,000+ tps on a 1T model wasn't a single breakthrough — it's deep model × system co-design between the MiMo and TileRT teams, all on general-purpose GPUs (no Cerebras-style wafer-scale, no Groq-style SRAM ASICs).
On the model side: FP4 quantization (smaller footprint, less memory traffic) + DFlash, our block-masked parallel speculative decoding that accepts far more tokens per verification. On the system side, TileRT tailors its compiler & kernels to exactly these techniques.
The result: a 1T model breaking 1,000 tps on a single, standard 8-GPU node.
🤗 Open weights (FP4 + DFlash checkpoint): huggingface.co/XiaomiMiMo/Mi…
Jun 8, 2026 · 2:37 PM UTC
10
49
580
42,893










