welcome to the lab. from the researchers at @scale_AI

You can’t improve agents if you can’t measure where and why they fail. At #LATechWeek, we're hosting a hands-on evals workshop led by our Head of General Agents @Bckenstler to break down how we structure evals, what we measure, and how we use the results to improve AI systems. Apply here to join us: scl.ai/LATechWeek
2
6
16
1,198
Scale Labs retweeted
Can research agents orchestrate agentic post-training with no expert in the loop? Watch the livestream and see it happen in real time. Will the agents compete or collaborate? Will they politely yield GPUs to one another like well-behaved coroutines, or race to grab and hoard as many as they can? Behold. Which frontier model makes the best post-training researcher? You'll get to weigh in at @COLM_conf 2026 in San Francisco. Co-organized by @bake_ai_hq, @Stanford, @NotreDame, @UW, and @scale_AI.
How much of post-training research can frontier models actually do on their own 🤔? We’re testing this with LIVE RSIArena, in collaboration with @Stanford, @NotreDame, @UW, and @scale_AI. ✨8 research agents. ✨Same 30B base model. ✨144 hours on a shared cluster of 64 RTX PRO 6000 Blackwell GPUs. 🎉 Model training model! Each agent gets 1,000 GPU-hours to choose its data, write training code, and run experiments. We freeze and independently evaluate the models they submit. Come watch the agents do research—and have their own “why did this work yesterday?” moments. If you work on post-training, agents, or evaluation, we’d love your feedback! RSIArena, by @bake_ai_hq 😆 Watch: 👇 We’ll also host a livestream event at the @COLM_conf. Come chat with us at booth #107!
4
2
16
2,578
Looking forward to attending COLM in SF with @ScaleAILabs next week ✈️ Recently I've been focused on user simulators to improve multi-turn evals and RL envs, would love to chat with anyone doing research in that space. Feel free to reach out if that's you!
4
1
11
432
Scale Labs retweeted
Glad to share that 3 papers I contributed to at @ScaleAILabs have been accepted to NeurIPS 2026, spanning post-training and agent evaluation 🎉 🔍 Reward Hacking in Rubric-Based Reinforcement Learning We conducted a systematic study of reward hacking in rubric-based RL, distinguishing failures caused by weak verifiers from failures in rubric design, and identified an interesting stopping criterion. arxiv.org/abs/2605.12474 🧠 Not Every Rubric Teaches Equally We introduce a policy-aware curriculum for rubric-based RL that emphasizes the criteria most useful for learning at each stage of training. arxiv.org/abs/2605.20164 🛠️ SWE Atlas We introduce a benchmark for evaluating coding agents beyond issue resolution, covering codebase Q&A, test writing, and refactoring. arxiv.org/abs/2605.08366 Really proud of the team and the work behind these. Looking forward to sharing more and catching up with everyone at NeurIPS in Sydney!
9
7
51
1,937
Scale Labs retweeted
Excited to share three @ScaleAILabs papers accepted to NeurIPS 2026 in the Evals and Datasets track! SWE Atlas: Evaluating coding agents beyond issue resolution, across codebase Q&A, test writing, and refactoring—with attention to both correctness and engineering quality. arxiv.org/abs/2605.08366v1 MCP-Atlas: Evaluating agents’ ability to discover and use tools across real MCP servers to complete realistic, multi-step tasks. arxiv.org/abs/2602.00933 Beyond Truthfulness (MASK): Separating honesty from accuracy by testing whether models contradict their own stated beliefs when pressured to lie. arxiv.org/abs/2503.03750v2 Agent evals and benchmarks are important as capabilities increase, safety becomes critical, and reliability remains an open question. Excited to continue this work with our customers, and grateful to all my co-authors and collaborators for the hard work behind these benchmarks!
9
4
29
1,024
Scale Labs retweeted
"MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers" has been accepted to NeurIPS 2026!! I had the privilege to be co-first author on the paper and to work with virtually all of the frontier labs to enable them to report numbers on our benchmark!
6
7
32
3,049
Scale Labs retweeted
🎉 Excited to share that “Reward Hacking in Rubric-Based Reinforcement Learning” has been accepted to #NeurIPS2026! Reward hacking in post-training is a particularly interesting and important problem to work on. We study how models can learn to game rubric-based rewards during post-training, satisfying the rubric without necessarily improving the broader response quality. We characterize these failures at the criterion level and show that stronger verifiers help, but don’t fully eliminate reward hacking. Link: arxiv.org/abs/2605.12474 Check out some of our recent post-training research projects at @ScaleAILabs: RGSD: arxiv.org/html/2606.12507 POW3R: arxiv.org/abs/2605.20164
4
8
28
933
Scale Labs retweeted
🎉 Excited that “Not Every Rubric Teaches Equally” is accepted to NeurIPS 2026! 🎉 POW3R turns static rubric rewards into policy-aware training signals, emphasizing the criteria that actually separate good and bad rollouts for the current state of the model. Link - arxiv.org/pdf/2605.20164 Our team at @ScaleAILabs has been cooking on some really exciting RL/post-training research lately. Check it out 👀 RGSD - arxiv.org/abs/2606.12507 Reward Hacking with Rubrics - arxiv.org/abs/2605.12474
4
7
51
2,385
We're releasing HLE-Diamond: a refined version of Humanity's Last Exam, built with @CAIS. A year of review and community feedback went into refining this subset to make it more reliable for measuring frontier models. Top model tested 60.6% overall. We expect HLE-Diamond to carry signal for the next 6-12 months.
4
15
96
7,153
Over a year ago, we released SWE-Bench Pro. We refreshed the benchmark today to improve its quality based on community feedback. Our update is meant to ensure the benchmark is clean and accurately measures model capabilities on SWE-Bench Pro tasks. In our update, we observe a ~20% performance drop on our held out private example set of 272 tasks compared to the public leaderboard. We further identify a HARD subset of 51 tasks that is discriminative of the frontier model performances from the rest of the pack. Refreshing benchmarks, and sharing what we learn along the way, does more to advance model evaluation than deprecating them outright. That's why we're sharing this updated version: it addresses issues we identified over time, and we want the broader community to benefit from those findings.
SWE-Bench Pro V2 is live. What’s new: 🧵
4
2
30
4,216
SWE-Bench Pro V2 is live. What’s new: 🧵
19
8
153
209,013
211 tasks were improved to have better dependency support for running OSS harnesses. We removed 89 invalid tasks, moving down to 642 tasks across 11 repos.
2
23
9,222
600+ proposals in 🤯 We're reaching out to top contributors to start building in our public repo. We’ve also welcomed @SchmidhuberAI, a pioneer of RSI, as a senior advisor to RSI Bench. New blog on our setup and verification pipeline: rsi-benchmark.com/blog/verif…
Launching rsi-benchmark.com: The work of AI R&D has always belonged to humans. For the first time, though, it no longer seems certain that it always will. Recursive self-improvement is within a line of sight. It may still be far, but it is close enough that we should start measuring it.
4
9
121
92,532
Scale Labs retweeted
AI safety doesn’t translate one-to-one. We partnered with the Korea AI Safety Institute to develop ROK-FORTRESS, our first benchmark together, testing how language and geopolitical context affect AI safety. Across 14 frontier models, we found that translation alone can miss meaningful differences in model behavior. scale.com/blog/korea-ai-safe…
5
6
44
15,761
New research from Scale Labs: SteerDuplex and SteerBench measure how well voice models follow instructions for how a response should sound and be delivered. +44.5 points on audio steering over the best open baseline. Pause barge-ins cut from 26.5% to 9%. Plus a sharp result on why timing rewards alone collapse into silence. ⬇️
Full-duplex voice models can listen and speak at the same time, but can they change how they speak when you ask them to? Introducing SteerDuplex + SteerBench: the first benchmark that measures steerable full-duplex dialogue. 65.1% audio-steering APR vs 20.6% for the best open baseline. 🧵
1
2
16
1,941
GPT-6 Astra is a step function on DrugDiscoveryBench at 68.7%, our benchmark for early-stage drug discovery computational workflows. Muse Spark 1.3 and Opus 5 follow. @TuXinming, @KavanaghOscar and I will talk about what it measures tomorrow at 10 am PT! Register below ⬇️
1
9
37
6,196