Incoming PhD @StanfordNLP | BS @PKU1898 | Building continually self-improving AI | Prev @DeepSeek_AI @AIatMeta | DeepSeek-V1/V2/VL/Prover ALE SPIRAL SPICE SPADE

South Korea
Continuous self-improvement needs an ever-expanding supply of training environments (goals). SPADE: one model self-plays the Environment Designer and the Reasoning Agent, writing executable, agentic environments that get harder as it improves. Environment scaling on its own. ♠️
20
127
703
178,695
Bo Liu (Benjamin Liu) retweeted
We release a benchmark that checks if retrievers (embedding models & ai agents) can find prior works that "inspired" a research paper, extending beyond simply finding relevant papers. In my point of view, having a good model on this benchmark can enable assessing novelty of papers or generating novel research ideas. Checkout @ohmyksh's post!
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
3
7
14
555
Bo Liu (Benjamin Liu) retweeted
In here is a video of The Bitter Lesson as a pretty-good country-music song. Enjoy. oneusefulthing.org/p/the-dot…
8
23
229
14,554
Bo Liu (Benjamin Liu) retweeted
To tackle new scientific problems, we often get inspirations from prior work. Can Deep Research agents and search help us to identify such “catalyst” papers buried in literature? Our new benchmark, built with hundreds of scientists, shows substantial room for improvements.
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
4
9
73
5,888
Bo Liu (Benjamin Liu) retweeted
Chief in AI 100 (2026): Yejin Choi, Professor, Stanford (MacArthur Fellow); Distinguished Scientist, NVIDIA. Profile: aibuildersnetwork.org/chief-… Meet leaders like these at the conference, Oct 14–16: aibuildersnetwork.org/confer… #ChiefInAI100 #aibgc2026
2
2
1,767
Research taste will matter more and more as AI starts doing research on its own. Part of taste is knowing which old papers to build on. ScholarCatalyst is a first step toward measuring that automatically :))
Great researchers have an uncanny ability to make connections that seem inevitable in hindsight, in places nobody else would have thought to look. Can we measure this ability? Introducing ScholarCatalyst: a far-from-saturated benchmark for finding what we call "catalyst papers"📚, labeled by 184 lead authors on 207 of their own recent projects. Paper: arxiv.org/abs/2610.02202 To make sustained progress on open-ended problems, I think agents need the sort of "research taste" that great researchers have. They need to make deep connections between earlier discoveries and problems those discoveries weren't intended to solve. ScholarCatalyst is a first step towards this goal. More details in the thread below🧵
2
7
40
2,292
Bo Liu (Benjamin Liu) retweeted
Great researchers have an uncanny ability to make connections that seem inevitable in hindsight, in places nobody else would have thought to look. Can we measure this ability? Introducing ScholarCatalyst: a far-from-saturated benchmark for finding what we call "catalyst papers"📚, labeled by 184 lead authors on 207 of their own recent projects. Paper: arxiv.org/abs/2610.02202 To make sustained progress on open-ended problems, I think agents need the sort of "research taste" that great researchers have. They need to make deep connections between earlier discoveries and problems those discoveries weren't intended to solve. ScholarCatalyst is a first step towards this goal. More details in the thread below🧵
16
49
310
21,600
Bo Liu (Benjamin Liu) retweeted
AI can increasingly make progress on open problems like Navier-Stokes. Yet defining a new problem still relies on human research taste. Introducing ScholarCatalyst, a benchmark built from AI researchers’ firsthand accounts of what inspired their work.
5
45
213
21,504
Bo Liu (Benjamin Liu) retweeted
FrontierPhysics update #3 We're wrapping up the v0.1 leaderboard. From initial runs on 56 real physics research tasks, the strongest frontier model passes only 22% of its runs. How we got a number we trust 🧵
2
4
14
2,199
Bo Liu (Benjamin Liu) retweeted
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
41
184
1,404
274,246
RT @jaseweston: ⏱️⚙️Introducing *AutoBenchmark* 📊🏁 - Creating benchmarks automatically - Benchmarking benchmark creation - Studying the ro…
64
4
RT @jaseweston: Claim: we've solved the AI slop problem (!) 💩🧹✨ Blog post: facebookresearch.github.io/R… 🧵1/5 Key idea: take *expert* human wr…
177
6
Bo Liu (Benjamin Liu) retweeted
🔥 Jev is on fire! 😎 So we ask Reef: /reefine add jev as a tool Reef then evolves the agent, and adds Jev into its harness. We use this evolved agent to hunt for papers related to "Self-evolving Agents" published in 2026 on arXiv. The agent found and screened 120 papers in 9 seconds! Come and use Reef to evolve your agent: github.com/Human-Agent-Socie…
7
26
61
104,358
Bo Liu (Benjamin Liu) retweeted
Reef just dropped v0.1.1 🚀 In the past week, Reef has: → 🌟 Crossed 5K GitHub stars, adding ~1.6K → 🔧 Contributions from 35 developers so far Try out v0.1.1 and tell us what you want your agent to keep learing. Checkout Github Repo at: github.com/Human-Agent-Socie…
1
34
135
38,129
Bo Liu (Benjamin Liu) retweeted
We built an all-synthetic simulation of a hospital that can be shared publicly without any issues. Our data quality is so high, physicians cannot tell the difference between real/synthetic patients. The data is fully verified -- perfect for RL.
Introducing Synthetic Hospital: an open, fully synthetic longitudinal EHR benchmark with verifiable ground truth! 1,268 patients, 5,602 encounters, zero PHI. Physicians could not reliably distinguish its charts from real ones. 📄 arxiv.org/abs/2609.30027 💻 github.com/sparkcpark/synthe… ✍️ sparkcpark.github.io/posts/f…
24
84
1,137
141,913
Bo Liu (Benjamin Liu) retweeted
This is cool! We’ll have grown up LLMs nurturing baby LLMs to learn by self-play in the crib in no time! More seriously, it’s a good demonstration of how general inductive biases can be an effective “Universal Grammar” for more general learning about the world.
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
4
13
116
13,440
Bo Liu (Benjamin Liu) retweeted
Your agent can now grow new abilities, just by asking. Reef now supports personalized harness evolution with /reefine: describe what you want your agent to do. Reef builds the change, checks it, and ships a new version. Examples: 💬 "> /reefine add a /chat mode for faster responses" 🔊 "> /reefine tell me out loud when you're done" ⚡ "> /reefine add jev as a tool": 120 papers screened in 9 s Works with any agent harness through a Reef adapter: Codex, Hermes, OpenCode, Pi and Terminus 2 today. Try it here: github.com/Human-Agent-Socie… #agent #harness #rsi #llm
9
39
119
164,609
Bo Liu (Benjamin Liu) retweeted
Introducing Synthetic Hospital: an open, fully synthetic longitudinal EHR benchmark with verifiable ground truth! 1,268 patients, 5,602 encounters, zero PHI. Physicians could not reliably distinguish its charts from real ones. 📄 arxiv.org/abs/2609.30027 💻 github.com/sparkcpark/synthe… ✍️ sparkcpark.github.io/posts/f…
53
146
1,303
298,306
DeepSeek will not announce «RSI», they'll just explain it in passing as section 6 of their paper on sandbox infrastructure
5
19
415
13,923
Bo Liu (Benjamin Liu) retweeted
Releasing our runtime dynamic compression framework that achieves 1.5-2.0 bit compression at high quality. This is integrated into the bitsandbytes2 library, which starts as a private beta today. Paper: timdettmers.com/papers/runti… Private beta signup: forms.gle/Pwtp8CVRULdoEnY79
20
48
294
31,248
Bo Liu (Benjamin Liu) retweeted
Starting tomorrow, our lab will hold an open-source week: 2 software frameworks, 4 papers, all building a coherent ecosystem. The theme: Frontier AI on Hardware You Own Blog post: timdettmers.com/2026/09/21/d… SoTA results in: Autocompaction Autonomous Research Model compression Test-time scaling Deep Research And a healthcare RL environment giving you a new level of complexity to train healthcare agents. The ecosystem that we will release is built to be as usable as possible. Agent sessions that run overnight and go on for tens of millions of tokens is made easy. Model compression of a model is automatic: you just get a good model that runs fast locally -- no expertise required. An autonomous research that works out of the box. Our ecosystem enables a new level of work that can be done locally.
28
75
554
37,197