Allen Institute for AI @allen_ai and @uwnlp. Machine Learning and Reinforcement Learning

Seattle, USA
CLMs let language models treat context as an editable file and learn to manage it themselves, rather than relying on hand-designed context-management rules. The idea behind CLMs goes beyond context management. I’m excited about the broader direction: agents that design and adapt their own harnesses in real time, like robots building and reshaping their own bodies and context as needed.
‼️The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA! Introducing 🩵Context Language Models (CLMs)🩵 - Natively manage their own context - Treat context as a file - Learn policies in CLM weights, no harness
4
210
Teng Xiao retweeted
We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have. Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to image-hosting sites as links that weren’t publicly listed. The images came from accounts that allowed their data to be used to improve our models, and after we disassociated the images from the accounts and ran them through a privacy filter. These cases occurred before the mitigations and safeguards we implemented and described in this blog post: openai.com/index/hugging-fac… We have successfully worked with the hosting providers to remove most of this content and are working to remove the rest. openai.com/hugging-face-inci…
860
688
5,181
2,249,368
Teng Xiao retweeted
GPT-6 Astra is state-of-the-art on: 🌟 Computer use 🌟 Browsing 🌟 Agentic coding 🌟 Cybersecurity 🌟 Science 🌟 Professional work
101
353
4,525
1,526,037
Teng Xiao retweeted
Pretty interesting rethinking paper on "discouraging" self-evolving loops. So the background is that most current self-evolving loops run their search directly on the test set. It kind of makes sense as harness search needs accurate, verifiable feedback to make grounded edits. However, that quietly turns self-evolving loops into a form of test-time scaling, which is exactly what this paper argues. Specifically, the paper points out that a loop that repeatedly evaluates and revises candidates against task feedback, then reports on those same tasks, is logically a test-time search procedure. So its gains should be measured against test-time scaling under matched feedback and inference budgets — otherwise you can't tell whether it discovered a better harness or just spent more compute. With that in mind, the paper runs four methods under the same budget: 1. parallel sampling — fixed harness, k independent trajectories per task, with a self-judge or unit tests picking the final answer 2. sequential refinement — fixed harness, k retries in a row. Each round summarizes the previous attempt into context and tries again (essentially prompt refinement) 3. harness evolution — the standard self-evolving loop. One shared harness, revised each round from feedback pooled across all tasks 4. harness scaling — the per-instance counterpart. Each task evolves its own harness The results are very interesting. Harness evolution doesn't beat plain parallel sampling, and without verifiable feedback it can even fall below single-attempt direct sampling with the initial harness. Its gains also show up at pass@5 but barely at pass@1, which implies that the improvement comes from taking multiple attempts, not from the harness getting better. And on a disjoint search/eval split the evolved harness transfers almost nothing to held-out tasks, which means the edits memorize task-specific fixes rather than distill reusable strategies. There is one caveat, which is that the "unified budget" only counts inference on the tasks, not the compute spent generating harnesses. But this flaw kind of favors harness evolution, and it already loses. So in my opinion this really shows that existing self-evolving loops might just be a different way of applying test-time scaling, rather than some new intelligence discovery. And we should focus more on making generalizable self-evolving loops work!
15
40
249
20,686
Submit your ai-generated work (and how you generated it) to our workshop! (automlr.com, openreview in the next few days) AI systems are increasingly producing scientifically interesting results, and attention and validation of results is becoming the bottleneck. Our workshop's format reflects that: selected reviewers act as discussants who present their own take on your work and open a discussion. We hope this shifts the focus to identifying what's worth other people's attention in a piece of research, alongside validating the core claims. As the world changes, so should our workshops and conference formats!
🚨🚨🚨 It feels like every other day a new breakthrough is attributed to an AI system. It is with this in mind that we are excited to announce our @NeurIPSConf 2026 Workshop on Autonomous ML Research! This workshop is for research where an AI agent performs end-to-end research or makes the decisive contribution to a machine learning discovery. Our goal is to make autonomous research meaningful in a way that strengthens the ML community without undermining the role of human researchers, so we introduce a couple of twists 🧵
1
4
15
2,545
In this work, we revisit how automatic harness evolution should be evaluated. Existing automatic harness evolution methods often search over harnesses using feedback from benchmark tasks and then report final performance on the same benchmark. This makes it difficult to tell whether the gains come from genuinely better and reusable harness design, or simply from spending more inference compute, receiving repeated task feedback, and adapting to the evaluation set. We compare harness evolution against simple test time scaling baselines under matched feedback and inference budgets. We also separately test whether evolved harnesses transfer to unseen tasks. On Terminal Bench 2.1, harness evolution does not consistently outperform parallel sampling or sequential refinement, either with or without unit test feedback. When the search and evaluation tasks are separated, the evolved harness provides only marginal improvements on held out tasks. Our takeaway is not that harness evolution is ineffective or unimportant. Rather, its benefits need to be assessed using fair experimental setups, strong baselines with comparable budgets, and benchmarks that are genuinely sensitive to harness design. We hope this work encourages more careful evaluation and helps identify settings where automatic harness evolution can produce real and transferable improvements.
Automatic harness evolution appears to be a promising path toward AI self-improvement, but we find that its gains still largely come from repeated sampling and show limited generalization. Blog post: yikee.github.io/harnessevolu… Code: github.com/rethinking-harnes…
1
17
7,210
Check out TMAX, a simple open RL recipe for terminal agents, led by @hamishivi and @yinn_oscar. A really nice step toward making terminal-agent training more open and reproducible. Strong results with a clean recipe, open data, models, code, and training artifacts.
Trained some terminal agents with friends! Introducing Tmax, open RL terminal agent models. Under default settings and shorter length (65k) token budgets, tmax outperforms prior open work on terminal use. We are releasing all data+weights+rollouts publically!
8
813
Recursive self-improvement (RSI) means the system improves the improvement mechanism itself. Each cycle produces not only a more capable system, but a system that is better at improving itself. RSI is always one order higher than the corresponding SI, because the recursion operates at the meta-level, on the rate of improvement itself. It could arrive faster than most institutions can adapt.
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…
468
Teng Xiao retweeted
LMs can learn from human labels, training data, and stronger teachers. But what happens when all of these run out🫪 when the model is already at the frontier and there is no stronger external source to learn from❓ In EvoLM, we extract the model's own evaluative knowledge into rubrics, and use them to improve its own generation🔁 This enables self-improvement with no external signals‼️
6
45
233
38,039
Congrats Huaisheng @huaiszhu on SDDLM being accepted to ICML 2026! 🎉 A personal milestone: my first paper as the last author. SDDLM uses a simple denoising objective for uniform-state diffusion LMs—lower cost, scaling to 1.1B, and strong generation quality. arxiv.org/abs/2510.22926
🚀 New work accepted at ICML 2026 Simple Denoising Diffusion Language Models (SDDLM) ⚡ Efficient training. Strong scaling. Our method reduces training cost while matching or surpassing prior methods, and scales effectively to 1.1B parameter models with strong performance.
1
2
14
1,609
My first paper since joining the ETH AI Center and the Apertus team: @AlexisLimozin discovered two bugs in DeepSpeed and OpenRLHF that lead to faulty SFT baselines often cited in mixed-policy RL for LLM reasoning methods. When applying the fixes, SFT-then-RL outperforms every mixed-policy approach that we tested. The bugs are: - DeepSpeed CPU-offload optimizer silently drops micro-batches in gradient accumulation (also hits TRL, OpenRLHF, Llama-Factory). - OpenRLHF misweights per-mini-batch losses.
We found two bugs in DeepSpeed and OpenRLHF that deflated SFT baselines in several recent mixed-policy RL papers. Once fixed, plain SFT-then-RL beats or matches every mixed-policy method we tested.
5
20
148
15,121
🚀 New work: Meta-Reinforcement Learning with Self-Reflection LLM agents shouldn't just solve problems. They should learn from their own attempts. Most current RL methods optimize single independent trajectories. Each attempt starts from scratch, with no mechanism to improve across attempts. But intelligent systems should get better after trying once. This raises a fundamental question: How do we train models to learn from their own attempts? We believe Meta-Reinforcement Learning may be a key paradigm for training future LLM agents, enabling models to adapt and improve across attempts and environments. In this work we introduce MR-Search, a training paradigm built around: 🧠 In-Context Meta-Reinforcement Learning 🪞 Self-Reflection 🔁 Learning to learn at test time 📄 Paper: arxiv.org/abs/2603.11327 💻 Code: github.com/tengxiao1/MR-Sear…
11
49
299
52,143
📊 Empirically, this simple idea works surprisingly well. Across multiple benchmarks, MR-Search significantly outperforms strong RL baselines, with 9–19% relative improvements. But the bigger takeaway isn't just the numbers. It's the training paradigm. 🔥 Even more interesting: Models trained with Meta-RL naturally benefit from additional reflection at test time. Performance keeps improving as we allow more reflection turns. This suggests Meta-RL enables true test-time adaptation: the ability for models to improve during inference. A broader perspective: Future LLM agents may not rely on one-shot reasoning. Instead they may operate in a loop: 🧠 attempt 🪞 reflect 🔁 adapt 📈 improve Training agents to learn from experience during inference may be one of the most important directions for building more capable AI systems. And Meta-Reinforcement Learning provides a natural training framework for this.
2
1
15
1,132
Huge thanks to my incredible collaborators: @1t4chiii, @hamishivi, @huaiszhu, @faeze_brh, @natolambert, @pdasigi, @nlpnoah and @HannaHajishirzi. We’re thrilled to see Meta-RL emerging as promising directions for training more adaptive LLM agents. 🚀
1
11
731
Teng Xiao retweeted
Small language models are not very helpful as judges, how about 🔄 backward inference—inferring the instruction given only the response, and using the similarity between the inferred and the original instructions as the reward signal? Introducing ⚙️FLIP, a reference-free and rubric-free reward modeling approach that boosts the RewardBench2 performance of 13 small language models by an average of 79.6%, and substantially outperforms LLM-as-a-Judge under test-time scaling via parallel sampling and GRPO training. 📄paper: arxiv.org/abs/2602.13551  🔗code: github.com/yikee/FLIP
12
52
251
29,007
Teng Xiao retweeted
🚀 New Paper Alert! 🧠 Simple Denoising Diffusion Language Models (SDDLMs) We simplify the complex ELBO objectives in Uniform-State Diffusion Models with a simple denoising loss, making training more scalable — while matching or surpassing baseline generation quality. (1/2)
1
3
2
689