How much of post-training research can frontier models actually do on their own 🤔?
We’re testing this with LIVE RSIArena, in collaboration with
@Stanford,
@NotreDame,
@UW, and
@scale_AI.
✨8 research agents. ✨Same 30B base model. ✨144 hours on a shared cluster of 64 RTX PRO 6000 Blackwell GPUs. 🎉 Model training model!
Each agent gets 1,000 GPU-hours to choose its data, write training code, and run experiments. We freeze and independently evaluate the models they submit.
Come watch the agents do research—and have their own “why did this work yesterday?” moments.
If you work on post-training, agents, or evaluation, we’d love your feedback!
RSIArena, by
@bake_ai_hq 😆
Watch: 👇
We’ll also host a livestream event at the
@COLM_conf. Come chat with us at booth #107!