The NLP group at the University of Washington.

Seattle, WA
KV Cache cry no more. Blog post by @RulinShao and @oscar_yinn
We recently released Context Language Models (CLMs), and got many questions on KV Cache. Our latest blog explains 1). how we measure cache-ware FLOPs 2). how to efficiently serve CLM wai-org.com/blog/clm/
1
12
1,995
UW NLP retweeted
Wrote a little blog with @RulinShao to visualize and explain some neat ways to reduce FLOPs usage for CLMs, check it out!!
We recently released Context Language Models (CLMs), and got many questions on KV Cache. Our latest blog explains 1). how we measure cache-ware FLOPs 2). how to efficiently serve CLM wai-org.com/blog/clm/
2
9
1,456
UW NLP retweeted
AI systems should be designed for AI. If prefix caching gets in the way of free context access, we can fix the caching. This is the idea behind our Suffix Cache Reuse (SCR) for CLR serving. Check out our blogpost for details (fun visuals included)! wai-org.com/blog/clm/ Maybe the next step is to give CLMs direct agency over the cache itself. I’m excited to see what more creative and adaptive cache-aware context management could look like. It's also a good time to follow our account @waiorg , a student org we created to share our work on agents. Currently covering: TMax, CLM.
7
25
158
8,228
UW NLP retweeted
‼️The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA! Introducing 🩵Context Language Models (CLMs)🩵 - Natively manage their own context - Treat context as a file - Learn policies in CLM weights, no harness
90
269
2,073
432,480
Image generation models CANNOT follow instructions (yet)🚨 We introduce Visual Verifiable Rewards (VVR) to solve it✅ VVRBench asks the model to "generate three yellow circles"🟡 Surprisingly, SOTA models, including GPT-Image-2.5-Sunburst, fail on complex VVR tasks. RL with VVR (RLVVR), and mixing VVR intro standard image post-training rewards, improves instruction following in natural prompts.
6
26
89
6,128
UW NLP retweeted
🎉 Excited to share that our paper “One Form to Transfer Them All” has been accepted to the #EMNLP2026 Main Conference! We study how pretraining representations shape cross-lingual transfer in multilingual LMs. 🧵
7
5
17
5,766
UW NLP retweeted
Super fun to see software world come to life!! We find codebases with large dependencies and build robust infra & envs where agents can meaningfully collaborate and work together. Demo: theagentorg.app
We share our early research on building Software World - a "GitHub" run by agents. We deploy agents for Python packages in a dependency chain, and each agent is tasked with collaborating with others and optimizing the package it owns.
1
4
989
UW NLP retweeted
Building an environment to develop and evaluate multi-agent systems at scale is challenging. Software World is our first step towards it: 📍We build a simulated GitHub ecosystem for agents that reflects the realistic needs of experienced software developers. 📍We construct robust, held-out, verifiable downstream applications that are less susceptible to reward hacking. Check out our blog post to learn more! theagentorg.app/
We share our early research on building Software World - a "GitHub" run by agents. We deploy agents for Python packages in a dependency chain, and each agent is tasked with collaborating with others and optimizing the package it owns.
2
8
91
8,686
UW NLP retweeted
Automatic harness evolution appears to be a promising path toward AI self-improvement, but we find that its gains still largely come from repeated sampling and show limited generalization. Blog post: yikee.github.io/harnessevolu… Code: github.com/rethinking-harnes…
8
61
310
64,997
UW NLP retweeted
I'm honored to receive the 2025 AAAI/ACM SIGAI Doctoral Dissertation Award! I recently spoke with SIGAI AI Matters about my dissertation, "Beyond Scaling: Frontiers of Retrieval-Augmented Language Models," and the research journey behind it: sigai.acm.org/main/ai-matter…
15
11
208
24,612
our very own @scottgeng00 on training olmo3 & dpo
New podcast/lecture combo -- a case study in the messy details of Olmo 3 post training & DPO with @scottgeng00. It's rare to make time for these discussions, but we cover: What it takes for a research idea to make it into a (near) frontier model. The messy side of DPO (usually data). Organizational challenges in multi-stage post-training recipes. Reflections on where research is heading today, and how people should think of DPO. DPO fundamentals (once more) Other topics. Scott is one of the great student's I've had the pleasure of working with for a few years! Chapters: 00:00 Introduction 04:32 Preference Tuning Basics 08:21 The DPO-Algorithm Zoo 17:06 Scaling Preference Data 21:55 Delta Learning Hypothesis 32:30: Applying it to Olmo 3 36:08 Challenge 1: Data Details 47:17 Challenge 2: Organization Management 51:45 Challenge 3: Scaling Issues 59:15 The Future of Academic Research Excited to keep making this content under the motivation of my book. Really plugging away now :)
1
5
5,396
Congrats to @hamishivi @yinn_oscar @RulinShao for their work on tmax 👐 credit to our data chef @yinn_oscar 🧑‍🍳
Scaling agentic RL environments: today we're publishing 365,000+ tasks for SWE, terminal, and search agents - 23 tasksets behind one API, one sandbox lifecycle, one command.
1
2
20
3,621
cool finding from @TengX6 and @yikewang_ on the evaluation of automatic harness evolution!
Automatic harness evolution appears to be a promising path toward AI self-improvement, but we find that its gains still largely come from repeated sampling and show limited generalization. Blog post: yikee.github.io/harnessevolu… Code: github.com/rethinking-harnes…
11
3,954
Check out this work led by @MA_Shafiei98! Models act fairer when the prompt looks like a fairness test⚠️ We call this **performative compliance** and introduce a method to reveal morally unsafe behaviors through verifiable demographic puzzles🧩
7
10
63
7,960
Most people think machine translation is solved, but there are so many other challenges even with frontier LLMs. Multilingual reasoning cascades need more context⚠️ This is the first paper I’ve mentored undergrads on and it’s both an exciting and scary process. I really learned a lot along the process, and so happy that I’m back to working on MT!! Truly a full circle moment when the paper came out🫶
Multilingual cascades are strong but lossy. We introduce a training-free fix: preserve the original question + context until final translation. Across 9 benchmarks, 3 models, and 285 languages, gains are strongest for cultural grounding, smaller models, and all resource levels.
1
3
31
4,026
UW NLP retweeted
Defended my PhD recently! Grateful to my advisors @LukeZettlemoyer and @nlpnoah, committee members Colin and @lucyluwang, and the many mentors, collaborators, and the @uwnlp community for all the support and friendship along the way ❤️ Slides: Break the language model monolith (drive.google.com/file/d/1na8…)
119
22
989
81,312
Excited to share our students are starting WAI @waiorg to do open agentic research. learn more at: wai-org.com check out their first work on open code agent recipe:
WAI is proud to introduce TMax: open terminal agents trained on an academic budget. Across different model sizes, TMax lies on the Pareto-optimal frontier of terminal-agent performance. Academic-scale RL can really produce highly competitive agents. Check out our Blog! wai-org.com/blog/tmax/
2
3
23
5,686
UW NLP retweeted
Introducing ><former Most transformers are rectangles◻️: every layer has the same width But is that optimal?🤔 We propose variable-width transformers that have different widths across layers, improving loss while cutting compute & KV cache size 🧵
62
71
828
348,204
What actually matters for MoEs? 🧐Great new work led by @margs_li
MoEs are everywhere, but the design space is confusing: total vs active experts? expert size? shared experts? routing? token dropping? We train >2000 MoE LMs 🫠 to investigate and bring you: 📄🔪🍰 Slicing and Dicing MoEs Tl;dr: it's all about expert size and count [1/9]
9
9,866
UW NLP retweeted
MoEs are everywhere, but the design space is confusing: total vs active experts? expert size? shared experts? routing? token dropping? We train >2000 MoE LMs 🫠 to investigate and bring you: 📄🔪🍰 Slicing and Dicing MoEs Tl;dr: it's all about expert size and count [1/9]
16
55
382
47,107