OpenBMB (Open Lab for Big Model Base) aims to build foundation models and systems towards AGI. Connect with us: discord.gg/7q3ry8Ny8K

Pinned Tweet
🚀 Meet MiniCPM5-2B, a 2B-parameter language model bringing high intelligence density to the edge, now open source! It ranks #1 among open-source models under 4B parameters on the @ArtificialAnlys Intelligence Index, with a score of 23. It also scores 20 on the Agentic Index, bringing an early form of general-purpose agent capability to the edge. Across 34 benchmarks, MiniCPM5-2B achieves an average score of 53.9, covering coding, math, long-context understanding, tool use, and agentic tasks. And this release goes beyond the model itself. We’re opening up the data, training recipes, and RL stack behind MiniCPM5-2B. 🤗 Hugging Face: huggingface.co/openbmb/MiniC… 💻 GitHub: github.com/OpenBMB/MiniCPM Modelscope: modelscope.cn/models/OpenBMB… Web: openbmb.cn/
189
319
2,300
657,858
A reward of “3” can mean two completely different things. Everyone thinks a response is mediocre — or half the people love it while the other half hate it. Most Reward Models cannot tell the difference. Introducing Diffusion Reward Models (DRM): instead of collapsing human preference into a single score, DRM learns the full reward distribution, preserving disagreement, uncertainty, and multiple plausible judgments. ✨ Paper:arxiv.org/abs/2609.33803 🤗 Models: huggingface.co/Teburile/DRM 💻 GitHub: github.com/thunlp/DRM Why it matters: Human disagreement is structured, not just noise. On datasets with repeated annotations, judgments often form separated or polarized patterns. More importantly, as human disagreement increases, DRM’s learned reward distribution becomes increasingly multimodal. The distribution is useful, not just descriptive. DRM can use distributional uncertainty to identify unstable reward decisions, and distribution-aware ranking improves Best-of-N selection beyond simply taking the mean reward. Reward Models get their own test-time scaling. Instead of only spending more compute on generating more responses, DRM can keep the response fixed and sample its reward distribution more times. More reward samples give a more reliable estimate — a new scaling axis unavailable to deterministic scalar RMs. And the gains survive RLHF. When used as the training-time reward, DRM improves downstream policy performance over scalar reward baselines, showing that the benefit is not limited to offline RM benchmarks. The takeaway: Reward modeling may lose something important when it compresses every human judgment into one number. “Everyone thinks this is average” and “people strongly disagree about this” should not look identical to a Reward Model. DRM makes that difference visible — and usable.
1
4
31
1,813
Post-training pipelines now use on-policy distillation (OPD) to hand a student the teacher's full next-token distribution at every prefix it visits: the student samples its own rollouts, and Qwen3, MiMo, GLM-5, DeepSeek-V4 and Kimi K3 all pair it with SFT and RL. Yet work on OPD has almost all stood on the algorithm side, treating the training set as given—so how much of OPD's gain does the data account for? Introducing One-Shot OPD, from @TsinghuaNLP (OpenBMB member) with the University of Chinese Academy of Sciences, Northeastern University, UIUC and Johns Hopkins University. It cuts the training set to one query, and the answer is that OPD is data-overfed but algorithm-starved. 1️⃣ One query, hundreds of steps. On math it goes from 59.1 to 68.5 by step 300, against 69.8 for full-data OPD—87% of its gain. It holds across code, instruction following and agentic tool use, and across Qwen, Llama and OLMo; a query the student never solves works about as well as one it always solves. 2️⃣ States, not questions. A prefix is a state where the teacher gives a target distribution, so 64 rollouts per step already yield tens of thousands of supervised positions. One query reaches 71.5% state coverage of what full-data OPD visits; 16 diverse queries reach 98.9% and match full data—extra queries buy new states, not new questions. 3️⃣Alignment slows, not the data. The student keeps improving, but each update absorbs less of what is left, which is why a run takes hundreds of steps rather than tens. This decline hardly depends on training-set size: on 1, 4, 16 and all 17k queries, alignment slowed at a similar pace. What limits a run is not how much data it gets, but how fast the student absorbs it. 📄 Paper: huggingface.co/papers/2609.0… 💻 Code: github.com/Thinking-Space/On… #AI #THUNLP #OpenBMB #LLM #PostTraining #Distillation #OpenSource
9
7
78
3,404
We hope to build an inference framework whose optimizations can be explained both when they work and when changing conditions make them fail. It is also the kind of engineering practice I hope more of us will undertake together in SGLang Omni.
Article

A Batch Size Debate: Performance Validation and Scheduling Trade-offs in SGLang Omni

In Revisiting CPU Resources as a First-Class Citizen in Speech Model Serving, we discussed a rather awkward problem: throughput from the same commit can change substantially with host contention.

4
2
4
2,270
MiniCPM-o 4.5 is now in SGLang Omni v0.1.7. More flexibility for developers to run and build with the model.
Hi everyone, today we released SGLang Omni v0.1.7. This release includes 75 merged PRs and welcomes 8 new contributors, with 8 first-time contributions. We added MiniCPM-o 4.5, NVIDIA PersonaPlex-7B, and OmniTyper powered by MLX streaming ASR, while further improving realtime and stateful Omni serving. 1.Performance: continued optimizations for Qwen3-TTS, Qwen3-Omni, CosyVoice3, MOSS-TTS, and AuK, covering Prefill CUDA Graph, speaker/reference encoding, kernel fusion, batching, and vocoder hot paths. 2.Serving: added Omni session lifecycle, the SGLang streaming session bridge, and a shared /v1/realtime WebSocket runtime, while further improving realtime ASR and streaming serving. 3.Models & hardware: added MiniCPM-o 4.5 multimodal input and speech output, plus PersonaPlex-7B offline speech-to-speech. MiniCPM-o and MiniMax-Music3 now support Intel XPU, with further MUSA support for Qwen3-TTS. 4.Runtime: improved breakable Prefill CUDA Graph, Talker / Code2Wav colocation, priority CUDA streams, scheduler admission, and profiling infrastructure to reduce host overhead and improve high-concurrency stability. github.com/sgl-project/sglan… github.com/sgl-project/sglan…
1
2
31
1,638
OpenBMB retweeted
🖥️ 上头了,AI Jarvis 这个开源桌宠会一直看着你的屏幕,你打游戏它在旁边搭话,你上网课它顺手把笔记记了 它跑的是 OpenBMB 的 MiniCPM-o 4.5 多模态模型,语言、视觉、音频三块全在你自己机器上算。安装包里连编译好的推理运行时都带着,不用自己去装 Python、CMake 或者 CUDA。 以前边看网课边记笔记是这样的:讲到关键那页要暂停、截图、切到笔记软件、打字,一节课来回切十几次,还总有几页没截着。AI Jarvis 的课程模式是它每秒看一帧画面、听最近一秒的系统声音,自己挑出关键画面,整理成知识点和画面说明,你从头到尾不用动手。 平时它就在桌面上待着,一个快捷键叫出来说话;打游戏的时候只在旁边发弹幕提示,不会替你操作。日常采集的原始画面和声音默认不长期保存,不想被看的时候有隐私模式能一键暂停感知。 目前只认 64 位的 Windows 10/11,第一次启动要联网下 6.32 GiB 模型,之后就是纯本地了,显卡不够会自动退回 CPU 跑。MIT 协议。 助理这东西,能自己看见你在干嘛,才算真的省事。 GitHub:github.com/LYiHub/pub-local-…
1
7
71
5,840
New article: how accurate are local models vs Jev, a cloud System One classifier? I ran 8 on a Mac, same JevBench questions. A 4B model matched Jev on short texts. On complex ones, all were 17+ points behind. rockyshikoku.medium.com/jev-… Video: minicpm5-2b classifying messages.
2
1
2
684
OpenBMB retweeted
Everyone is posting 3 and 4 Spark clusters this week. Here is what ONE DGX Spark does for one user, all measured on my box: MiniCPM5-2B, OpenBMB's drafter on: 100.8 tok/s Qwen3.6-35B-A3B: 89.7 Ling-3.0-flash: 69.5 Qwen3.8-Flash-Next: 55.5 DeepSeek-V4-Flash: 47.9 in EXL3 Gemma-4-E2B: 39.2 Qwen3.8-27B: 33.4 EXL3 + DFlash2 Tinfield-1 at 2-bit: 30.7 One model at a time, 256 tokens in, 256 out
22
13
87
5,108
OpenBMB retweeted
2B 参数的小模型,现在也开始能干活了? OpenBMB 的 MiniCPM5-2B,主打本地小硬件运行。有人把它接进 CI 修复 Agent,丢了一个“结算重试导致运费重复计算”的 Bug 给它。 模型自己复现问题、翻代码、定位原因,最后直接打补丁,18 个测试全部通过,而且全程本地运行。 几个亮点: ① 编程和工具调用能力不错 ② 2B 参数,对本地设备更友好 ③ 权重和训练数据都开源 小模型不一定只能聊天,拿来跑本地 Agent 也开始有点意思了。 🔗 huggingface.co/openbmb/MiniC…
13
2
18
3,191
⚡ Making MiniCPM5-2B even more lightweight for local inference! A community-built EXL3 4-bit quantization of MiniCPM5-2B brings the quantized model weights down to just 1.61 GB. ✨ Highlights: ⚡ 4.0 bpw EXL3 quantization for a smaller model footprint 🚀 ~68–70 tokens/s verified on an NVIDIA Tesla T4 💾 Just 1.61 GB for the quantized model weights 🛠️ Supports local inference with ExLlamaV3 and TabbyAPI A great community contribution showing how MiniCPM5-2B can be optimized for more lightweight and accessible local AI deployments. 🤗 Model: huggingface.co/ewin-reg/Mini… 🤗 Base model: huggingface.co/openbmb/MiniC…
8
19
155
6,475
OpenBMB retweeted
The first demo showed 32 local agents generating text at once. This time I gave them tools and put GPT-6 Astra in charge. MiniCPM5-2B handled the workers on my DGX Spark. Astra coordinated through Codex CLI. Their job: restore 32 services in a simulated city blackout using shared, limited repair supplies. Then a second failure closed a bridge and knocked six restored services offline. You can see tools reject actions, workers get blocked, and Astra send more fuel before assigning another round of repairs. It chose to leave the bridge closed and use the alternate route. 232 actual tool calls. 19 rejected actions. All 32 services verified online in 51.5 seconds. This is the setup I wanted to test: a stronger model coordinating small local models that can act and check their work. Full live recording below. 30 fps, no speedup.
6
2
28
2,054
OpenBMB retweeted
open source is on fire this week! for some weeks, i've been playing with very small LLM models in my free time (gemma e4b, qwen 4b, etc.) just to see what “small” stuff these LLMs can do. I got access to this 2 billion dense param model MiniCPM5-2B, and this thing is a beast for its size. 128k context window and better than every model in the same size range running fully locally on my macbook pro m5 (via llama.cpp)
17
3
188
9,302
OpenBMB retweeted
Introducing JustRL II 🚀 Building on JustRL (x.lingyaoai.com/HBX_hbx/status/1988474…), we took a closer look at how GRPO behaves in long-CoT RL (128k). The group-mean baseline is a great fit for short traces, but over tens of thousands of tokens it gives a coarse, response-level signal. JustRL II keeps GRPO's group structure and adds a critic for token-level credit assignment. Just as simple, keeps improving where GRPO levels off: AIME25 61→81 on a 2B model. The same recipe powers the RL stage of MiniCPM5-2B (huggingface.co/openbmb/MiniC…), making it SOTA among models under 4B. Data + checkpoints are open. Code lands this week. 📖 panhaoxuan.notion.site/justr…
✨What if the simplest RL recipe is all you need? Introducing JustRL: new SOTA among 1.5B reasoning models with 2× less compute. Stable improvement over 4,000+ steps. No multi-stage pipelines. No dynamic schedules. Just simple RL at scale. 📄 Blog: relieved-cafe-fe1.notion.sit…
5
41
286
26,833
OpenBMB retweeted
GGUF quantization may be getting a size slider of sorts with a new FIT-GGUF quants! Normally, when downloading a local model, you choose a quant level Q3 → Q4 → Q5 → Q6 …and hope it fits your RAM/VRAM budget. FIT-GGUF flips that around. Give it a target size or fidelity tier, and it can mix precision tensor-by-tensor to find a GGUF that meets that target. OpenBMB highlighted it on MiniCPM5-2B: 📦 Quality — 1.46 GiB ⚖️ Balanced — 1.28 GiB 🗜️ Compact — 1.21 GiB 🤏 Mini — 1.14 GiB Those labels represent progressively looser fidelity to the BF16 model, measured primarily with KL divergence and not different skills like coding vs. reasoning. And they run in normal llama.cpp / LM Studio. ⚠️ FIT does not claim its tensor allocation is universally optimal yet. 🔗 Link in ALT
7
10
58
4,720
🚀 FIT-GGUF brings controllable-size mixed-precision quantization to MiniCPM5-2B Developer @Scorp1o_117 used FIT-GGUF to build MiniCPM5-2B GGUF variants around specific size and quality targets. Instead of choosing a fixed quantization preset, you can set a target file size or fidelity tier, and FIT-GGUF automatically decides how much precision to allocate to different tensors—then predicts, generates, and verifies the final GGUF. ✨ What’s included 🧠 Tensor-level mixed-precision quantization 📦 Four MiniCPM5-2B builds from ~1.14 GiB to ~1.46 GiB 🎯 Quality / Balanced / Compact / Mini presets 📊 KL Divergence and Same-top evaluation ✅ Generated file sizes matched the predicted targets A nice example of how MiniCPM5-2B can be tuned for different memory and deployment constraints, without being locked into a single Q4/Q5-style quantization preset. Check out FIT-GGUF and try building a MiniCPM5-2B variant that fits your own device budget. 🤗Model: huggingface.co/SC117/MiniCPM… huggingface.co/openbmb/MiniC…
3
1
40
1,875
OpenBMB retweeted
First shot on @OpenBMB based Augury - weeds as indicators app. Time to iron out the bugs 🦠👌 Offline AI, sovereign based, farmer based AI.
2
1
4
364
🎙️ Real-time interpretation, powered locally by VoxCPM2. Developer @HenryZ30734018 built VoxWeft, an open-source simultaneous interpretation system for Apple Silicon. It uses an MLX implementation of VoxCPM2 to turn live speech into translated speech on-device, keeping audio private and responsive. ✨ Highlights: ⚡ VoxCPM2 streams first audio in ~170 ms on an M5 MacBook 🌍 Generates speech across 30 languages, supporting direct language-pair interpretation without a pivot 🗣️ Clones a target voice from ~5 seconds of reference audio for a consistent interpreted voice 💻 Runs in 4-bit quantization on MLX, making low-latency local speech generation practical on Apple Silicon VoxWeft shows how VoxCPM2 can become the speech layer of a full real-time application: not just producing audio, but enabling private, multilingual interaction that stays on-device. Try VoxCPM2 and see what you can build with it! 🔗GitHub:github.com/HenryZ838978/VoxW… 🔗Devlog:github.com/HenryZ838978/VoxW… 🤗 VoxCPM2:huggingface.co/openbmb/VoxCP…
3
2
47
2,128
What can a 2B open-source model actually do on your phone? MiniCPM5-2B from OpenBMB is 2B params, #1 under 4B on the Artificial Analysis Intelligence Index, and tops their Agentic Index outright, 20 vs 9 for the next small model (official). That's the Densing Law in action: model capability density doubles roughly every 3.5 months, and this one fits on an iPhone. The thing I want from a model that size is an iMessage agent: reads my texts, makes reminders and events, looks things up, replies for me, nothing leaves the phone. So I tested it. Six tools, eight real-world scenarios, three runs each. My results, not official ones. Worked: "remind me to send Jordan the invoice tomorrow at 9am" → correct reminder, 6/6. Delta confirmation text → clean JSON of flight, times, seat, 6/6, ~1.5s. "is the 6 train running Saturday?" → searched, read the result, replied correctly: take the 4 or D. "book it and tell her yes" → calendar event plus reply to Maya, 2/3. Didn't: "dinner thursday 7" became Sept 14, not the 10th, in most runs. Once it reasoned the right date and then wrote the wrong one anyway. Decided Lucali is in San Francisco. It's in Brooklyn. Thinking mode ran away: 4,096 tokens on a Swift date parser, no code, three times. HuggingfaceRepo:huggingface.openbmb.cn/model…
19
13
38
157,742
OpenBMB MiniCPM-V4.6 é um dos sinais mais claros de que a IA multimodal está ficando pronta para edge AI. ~1.3B parâmetros. Roda localmente em hardware de consumo. E ainda lida muito bem com OCR, dashboards, gráficos e compreensão de documentos. Testei no iPhone + Mac. 🧵
Paid partnership (ad)
12
19
38
57,534
Thanks so much for sharing this! Really cool to see MiniCPM5-2B being used in a practical multi-agent workflow like this — especially with the workers actually handling matching, short payments, duplicate references, and disputes through tool calls. Appreciate all the work you put into testing and documenting this. Such a nice case for the community 🙌
Next test on Spark 1: 32 synthetic invoices, with purchase orders and payment records. GPT-6 Astra coordinates the MiniCPM5-2B workers. They sort out matches, short payments, duplicate references and price disputes, then write the results into a case ledger. All 32 verified in 67.8 seconds. Eight in each category, with 232 executed tool calls. The video is real time. This demo doesn’t move money.
3
23
1,722
我们期待一起建设的推理框架,能够说清楚优化为什么有效,条件改变后又为什么失效。把一个反常结果追到可以解释、可以复现,再把这些认识写进代码和测试里,本身就是很扎实的系统工作,也正是我希望更多朋友在 SGLang Omni 中共同完成的工程训练。
Article

从一次 Batch Size 争论,思考 SGLang Omni 的性能验证与调度取舍

在《重新审视 CPU 资源作为语音模型 Serving 过程的一等公民》一文中,我们讨论过一个颇为尴尬的问题:同一个 commit

5
4
39
6,398