Run LLMs fast at any scale 🔗 github.com/sgl-project/sglan… Join our community slack.sglang.io For AI tech blogs & deep-dives 👉 @lmsysorg

Palo Alto
Pinned Tweet
Announcing the first SGLang Summit: November 12–13 at Fort Mason, San Francisco. Join us for two days of talks, discussions, and deep dives into frontier AI infrastructure, silicon, models, and applications. Register: sglang.io/summit
16
45
201
1,596,410
The SGLang community was absolutely buzzing at AI Infra Summit! Highlights from our sessions last week in Santa Clara: 💡 Alex Nails (@alxnails, SGLang core contributor) with Vasanthi Jagatha of @SamsungSemiUS on SGLang HiCache + Samsung Cognos for the KV cache bottleneck 💡 Sundara Raman Ramachandran of @LinkedIn on running latency-critical ranking on SGLang at global scale. Thanks to everyone who stopped by our booth! Excited to see you at our upcoming events: sglang.io/events #SGLang #LLMInference #AIInfra
1
1
26
2,136
SGLang v0.5.21 landed! Native decisions API is here 🎉 Some of our favorite updates: - Decisions API turns an LLM/VLM into a low-latency classifier and scorer - /v1/score can now rerank search or RAG results in one go - PD instances can switch between prefill and decode with no restart needed - DeepSeek-V4.1 Flash gets 22% faster first token on long prompts - Kimi K3 gets 20.6% higher prefill throughput in PD serving - GLM-5.3-Flash now runs on AMD MI355X with FP8 / MXFP4 MoE and MTP - You can now run MiniMax H3 inside @ComfyUI with SGLang-Diffusion backend New models include DeepSeek-V4.1 Flash, GigaChat 3.5, MiMo-V2.6, Ling-3.0-flash-VL, IQuest-Q1, Qwen-Image 2.1, FLUX 3 Action, and more. Full release notes👇
16
22
89
8,007
Last release:
SGLang v0.5.20 landed! Welcome @intel XPU to join standard SGLang releases 🎉 Some of our favorite updates: - RL sampling masks make rollouts more reliable, with up to 52% faster decode - Unified Radix Tree adds SWA branching-point caching: ~20pt higher cache hit rate, ~1/3 lower TTFT - DSpark now supports PD + DCP for long-context serving - SGLang Simulator brings scheduler & cache experiments to CPU - ROCm model loading is up to 12.5× faster - SGLang-Diffusion gets up to ~38% lower E2E latency New models include GLM-5.3-Flash, Qwen3.8-Flash-Next, K2 Horizon, Hy4-Preview, FastH3, VDN-H3, and more. Full release notes👇
260
Blog post about the update to scoring API:
🚀 New blog: Scaling JEV-like decision models with SGLang Decision models need a score, not prose. Classification, ranking, and agent action selection all ask the same thing: which option wins? Open-Jev, for example, scores each candidate separately with a Yes/No prompt. Serving this well raises two issues: Generate + top-k logprobs can drop the label you need, and the shared context can be recomputed for every candidate. SGLang addresses both: - /v1/score returns scores for the exact labels you request (Yes/No, A/B/C) - Multi-item scoring (MIS) computes the shared context once and keeps each candidate isolated - MIS latency stays nearly flat from 2 to 16 candidates, with 16-candidate p95 on Qwen3-8B dropping from 54.1 ms (Generate) to 20.6 ms (MIS) - MIS p95 stays under ~100 ms as load rises on Qwen3-0.6B, vs. seconds for Generate and SIS Huge thanks to the @LinkedIn team for contributing! Benchmarks and launch commands in the blog 👇
1
122
Decisions API demo:
We turned Qwen3.8-27B into a multimodal decision model. It beat Pokémon FireRed’s elite four and champion with sub-100 ms decisions from live game state. With SGLang’s native /v1/decisions, you can now turn LLMs and VLMs into classification and scoring models. We also added /v1/systemone so Jev-like open models can work with the TypeSafe SDK.
2
192
SGLang is headed to @modal Runtime tomorrow! 11:05am — @BanghuaZ on building frontier AI infra with SGLang + Miles 8:30am–6:30pm — Come meet the SGLang engineers. Bring your inference, serving, and RL questions 💻 See you in SF! 👇
Runtime is this week at The Midway in San Francisco. Thanks to our partners for helping us put it together: @nvidia, @pydantic, @radixark, @sgl_project, @temporalio, and @inferact, @vllm_project 💚 Join us to learn from the engineers behind these tools.
1
2
22
3,975
Announcing the first SGLang Summit: November 12–13 at Fort Mason, San Francisco. Join us for two days of talks, discussions, and deep dives into frontier AI infrastructure, silicon, models, and applications. Register: sglang.io/summit
16
45
201
1,596,410
We turned Qwen3.8-27B into a multimodal decision model. It beat Pokémon FireRed’s elite four and champion with sub-100 ms decisions from live game state. With SGLang’s native /v1/decisions, you can now turn LLMs and VLMs into classification and scoring models. We also added /v1/systemone so Jev-like open models can work with the TypeSafe SDK.
97
348
3,281
249,543
One of the most exciting parts of Pokémon is type matchups and knowing when to attack, switch, or heal. Qwen3.8-27B is a surprisingly good Pokémon player. It cleared the Elite Four and Champion in one go. Full run playback👇
7
5
111
9,732
SGLang has Day-0 support for IQuest-Q1 from @IQuest_research. IQuest-Q1 is an open-source sparse MoE model for CLI agents. It's built for coding, software engineering, and long-horizon agentic tasks. - 320B total params, 15B active per token - Trained with multi-harness RL and multi-teacher on-policy distillation (MOPD) Serve it with SGLang today:
Today we're releasing IQuest-Q1 and opening the model weights. 320B total. 15B active. Built for code, software engineering, and complex agentic tasks. Technical report/HuggingFace/GitHub — available now.
4
4
28
5,081
SGLang is a perfect fit for Jev-style inference, with efficient prefix caching and fast, stable structured generation. Here are some of our fav Jev projects 🧡 1/6 openjev-sglang by @ekzhang1 64 decisions in under one second, powered by Qwen3.6-35B-A3B and a single shared prefill reused through SGLang’s Radix Cache. x.lingyaoai.com/ekzhang1/status/210065…
Inspired by @typesafeai , here is a Jev-compatible public API to play with It runs a comparable open model (Qwen3.6-35B-A3B), and just uses SGLang radix cache to preserve the prefill reuse / really fast parallel systemone generation - 64 tasks in <1s. github.com/ekzhang/openjev-s…
7
9
48
16,491
3/6 openjev by @justALEXWORTEGA A family of Jev-style models spanning 0.8B, 2B, 4B, and 35B-A3B for reranking, grading, guardrails, and real-time game decisions, with SGLang serving support. x.lingyaoai.com/justALEXWORTEGA/status…
Typesafe: pnewed 💨 Jev: liberated 🫡 I trained an MLP on top of qwen 4b and it works literally like JEV huggingface.co/AlexWortega/o…
2
3
819
5/6 diffusion-jev-sglang by @hangzhi_wei Use @googlegemma DiffusionGemma with SGLang to build “Jev with eyes 🐈” for text and image decisions, with no task-specific fine-tuning. x.lingyaoai.com/hangzhi_wei/status/210…
My favorite thing about DiffusionGemma as #Jev is that it has eyes. I served it with #SGLang, and it recognized my terrible cat drawing 🐱 #doodle It gets close to Jev’s accuracy, with lower latency in my local setup. Everything works out of the box.
1
1
1
390