Don't pick between better prompts and a better model: improving both in turns lifted a science agent from 42.2% to 73.3% accuracy. Researchers correct AI agents all the time, but fixing an answer in chat doesn't make the agent better at the next task. Correcting an AI agent in chat, the fix usually dies with the conversation. ScienceBuddy, an AI assistant for scientists, turns that feedback into scored test tasks. Then it takes turns: rewrite the agent's prompts and skills, retrain the model, and repeat. With a small 4B model on biology tasks, prompt and skill changes alone raised accuracy from 31.1% to 51.1%. Retraining alone also helped, raising the share of problems solved within 4 tries from 48.3% to 67.8%. If you build agents, save every user correction as a test, and keep upgrading your prompts and your model in turns. – arxiv. org/abs/2609.17523v1 Title: "ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents"

Oct 2, 2026 · 8:47 AM UTC

8
12
66
5,582
Sort replies: Relevant Recent Liked
Replying to @rohanpaul_ai
Thanks buddy… 😀
1
2
33
😃🫡
1
87
Replying to @rohanpaul_ai
I keep hitting this building a reading helper for language learners. Someone corrects an explanation once, and the next week it makes the same mistake for someone else. Making that one correction stick turned out harder than writing the explanation.
1
19
Replying to @rohanpaul_ai
Agents may need something closer to an evolutionary loop: failure becomes a benchmark, the benchmark changes the system, and the improved system generates new failures to learn from.
23
Replying to @rohanpaul_ai
prompt and skill changes alone taking a 4b from 31% to 51% is the part I would read twice, most people jump straight to retraining
2
30
Replying to @rohanpaul_ai
A 4B model going from 42% to 73% on saved corrections is a better self-improvement argument than most AGI roadships out there.
20