Hi everyone, today we released SGLang Omni v0.1.7. This release includes 75 merged PRs and welcomes 8 new contributors, with 8 first-time contributions. We added MiniCPM-o 4.5, NVIDIA PersonaPlex-7B, and OmniTyper powered by MLX streaming ASR, while further improving realtime and stateful Omni serving.
1.Performance: continued optimizations for Qwen3-TTS, Qwen3-Omni, CosyVoice3, MOSS-TTS, and AuK, covering Prefill CUDA Graph, speaker/reference encoding, kernel fusion, batching, and vocoder hot paths.
2.Serving: added Omni session lifecycle, the SGLang streaming session bridge, and a shared /v1/realtime WebSocket runtime, while further improving realtime ASR and streaming serving.
3.Models & hardware: added MiniCPM-o 4.5 multimodal input and speech output, plus PersonaPlex-7B offline speech-to-speech. MiniCPM-o and MiniMax-Music3 now support Intel XPU, with further MUSA support for Qwen3-TTS.
4.Runtime: improved breakable Prefill CUDA Graph, Talker / Code2Wav colocation, priority CUDA streams, scheduler admission, and profiling infrastructure to reduce host overhead and improve high-concurrency stability.
github.com/sgl-project/sglan…
github.com/sgl-project/sglan…