Pinned Tweet
Wrote a short blog about the "shape" of language models, and the tradeoffs they may present in the future. I genuinely think it's a valuable research direction to start thinking about now, especially w.r.t. harness design. alexzhang13.github.io/blog/2…
42
209
1,973
210,526
alex zhang retweeted
Yep! Claude Science has a python repl with a host bridge that gives it access to harness primitives so it can do lots of things programmatically. MCP tool calling, skill creation, agent profile creation, LLM calling, delegation, and more :) The python repl is also perfect for doing computational science since bulky datasets stay in memory as agents iterate on their analysis!
5
11
123
10,062
alex zhang retweeted
late to reading this, but FUCK YES to playing with more shapes for intelligence! please play around with it more! append-only chat is not the final interface!
Wrote a short blog about the "shape" of language models, and the tradeoffs they may present in the future. I genuinely think it's a valuable research direction to start thinking about now, especially w.r.t. harness design. alexzhang13.github.io/blog/2…
20
20
337
22,137
alex zhang retweeted
SFT is not dead! 🥳 We found a way to make SFT rival current prevailing posttraining methods, often generalizing better and forgetting less than RL and OPSD. 🤯 Following our prior work on reasoning with sampling, we now introduce sampling to the posttraining stack. 1/n
60
212
1,844
256,402
local AI is the way! Congrats on the launch :)
Introducing Underdog, your Private Personal AI on devices you already own Today we're announcing our backing from @a16z @khoslaventures @HummingbirdVC Anthology (@AnthropicAI @MenloVentures) @patrickc @naval @rauchg @polynoamial @Thom_Wolf @Mascobot @tszzl @OfficialLoganK and other top AI leaders Our mission is to provide free, capable, reliable and private AI to billions of people. Underdog’s Law: today’s frontier intelligence reaches your devices in six months. We’re starting with fast, capable AI that runs on consumer hardware. By co-designing models and inference engines, we’re pushing the frontier of capability, speed, power and data efficiency. Our research also spans agentic commerce, confidential inference, and how AI will reshape the internet economy. Privacy and capability no longer needs to be a tradeoff. If you believe in this future, join us. Time to build.
1
3
72
7,732
alex zhang retweeted
Recursive Language Models: Claude Code, Agent Swarms, Big Research Bets, & the $40M AI Problem Solver latent.space/p/rlm MIT’s @a1zhang explains why Claude Code, Codex, and Pi are basically the same, how RLMs use code, context offloading, and recursive subagents to generalize across tasks, why one expert can sometimes replace a trillion-token brute-force search, what OpenAI’s 10,000-agent, 130B-output-token experiment reveals about the future of language models, and why academia’s biggest advantage is the freedom to take weird, ambitious research bets.
14
5
56
6,414
Is Jev secretly learning a Value function? A model can always choose “A” without being certain that “A” is correct. That distinction matters when an agent must decide whether to act fast or switch to deeper reasoning. We spent some time exploring: 1. How Jev fits into System 1/2 decision-making and why calibration matters for knowing when to switch 2. Why maximising expected reward in RL does not guarantee calibrated probabilities. 3. How predicting eventual task success in agentic tasks connects to a familiar RL quantity: the Q-function. Could value prediction help build a calibrated System 1—and could that be part of the story behind Jev? Link: souradip-chakraborty.github.… #Jev #RLCD #System1
3
23
111
9,391
alex zhang retweeted
Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do. SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
67
87
828
587,899
few thoughts, I've been sent this paper many times today, and it's cool! 1. it's a v clever idea, and I'd say even more extreme than RLMs on the spectrum of ReAct-style vs. pure context offloading (so no they're not the same) 2. im hopeful to see harness designs that use this principle, and they somewhat go hand-in-hand with model + RLM progress as well 3. the KV issues scare me admittedly, and the proposed fix is a bit hacky despite it seeming to work well for their results. that being said, I suspect different model shapes won't have this issue, so it's fixable! another + for this direction
‼️The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA! Introducing 🩵Context Language Models (CLMs)🩵 - Natively manage their own context - Treat context as a file - Learn policies in CLM weights, no harness
20
41
691
52,271
alex zhang retweeted
Language Model "Shape" Most AI agents force the harness around a decoder-only Transformer. But what if we change the model’s input/output shape to fit the agent? This blog by @a1zhang suggests using recurrent memory for old history while dense attention handles recent context, reducing the need for manual compaction. Read more about it: alexzhang13.github.io/blog/2…
4
20
187
8,147
Many libraries nowadays are agent-written, so a natural question is whether they can design & use these libraries effectively. It’s important to track this behavior, especially as agents continue to be used for coding. Awesome work by @GOrlanski that I got to play a role in!
Agents will soon replace humans as the main users and designers of libraries. A good library will let future agents write correct programs with less code. Introducing LibraryDesignBench: one agent designs a library, and other agents write programs with it.
2
5
93
7,993
The best club at MIT is starting up again! We'll be hosting talks with speakers on various interesting topics. This Thursday will be Alok, who wrote the famous blog / video on group theory and positional encodings! Sign up for the mailing list: scale-ml.org
Scale ML is resuming this Fall with Alok from Jane Street! We will go over how to do data mixture experiments at scale! 🔥 Sign up for the zoom link via the website.
3
30
536
43,491
alex zhang retweeted
First author of this paper here. Just want to give a special shoutout to @a1zhang’s RLM work since the general takeaway is very similar. The RLM work showed that a minimal code-mode harness can handle large inputs without needing external tools our harness components. We show that the natural generalization that makes *all* inputs (incl. user prompt instructions) and the REPL history itself variables in the REPL lets you do even more without external tools/harness components, such as very long-horizon tasks and self-improvement!
Keep your agent harness minimal, folks. This paper from MIT CSAIL shows why. (bookmark it) While being very minimal it looks effective and promising for building self-improving agents. They present JAZ, an agent framework with one primitive, invoke. The LLM writes code, can call invoke recursively, and sees all of its inputs and history as variables in the code environment. With prompting only and no memory system, it beats Letta (MemGPT) by 8% at half the cost on the recall-heavy part of StuLife. On AppWorld it beats ACE, a self-improvement method, by 4% at lower cost. Memory and self-improvement are usually built as separate subsystems. Here both are code the agent writes inside the loop, and hooks handle the constraints and monitoring you want to enforce. Paper: arxiv.org/abs/2609.26891 Explore Paper: academy.dair.ai/papers/harne…
12
15
174
17,286
alex zhang retweeted
Keep your agent harness minimal, folks. This paper from MIT CSAIL shows why. (bookmark it) While being very minimal it looks effective and promising for building self-improving agents. They present JAZ, an agent framework with one primitive, invoke. The LLM writes code, can call invoke recursively, and sees all of its inputs and history as variables in the code environment. With prompting only and no memory system, it beats Letta (MemGPT) by 8% at half the cost on the recall-heavy part of StuLife. On AppWorld it beats ACE, a self-improvement method, by 4% at lower cost. Memory and self-improvement are usually built as separate subsystems. Here both are code the agent writes inside the loop, and hooks handle the constraints and monitoring you want to enforce. Paper: arxiv.org/abs/2609.26891 Explore Paper: academy.dair.ai/papers/harne…
79
66
516
49,268
Wrote a short blog about the "shape" of language models, and the tradeoffs they may present in the future. I genuinely think it's a valuable research direction to start thinking about now, especially w.r.t. harness design. alexzhang13.github.io/blog/2…
42
209
1,973
210,526
I had this thought around the unnatural shape of language models (i.e., the autoregressive, decoder-only design of LMs) in RLMs, but was very excited with the release of @CompleteSkeptic / @typesafeai's Jev, which empirically offered an example of one such useful tradeoff. I think what once was something you'd never think to do, now can be done because of the availability of powerful language models that can be distilled to new model shapes.
2
2
63
8,808
I have some ongoing work here that is still very much experimental, but I'm sure there's a lot of other cool ideas out there worth thinking about as well. Overall an exciting direction and my DMs are open for those who are looking in it :)
2
1
37
5,531
Using open benchmarks has gotten very frustrating because it's very hard to disentangle what the model already knows / is actually bottlenecked by. E.g. when testing RLMs on "new" long benchmarks, some models will just magically regex for the right offloaded information and grab the solution... This is roughly how I've interpreted the HarnessTax paper as well, where performance seems to basically be a function of only the model capability and not the harness, but in practice we observe very different behavior / preferences with different harnesses. Model capability dominates on open benchmarks specifically. Not really sure what this implies because newer benchmarks constantly get swallowed, but I hope someone comes up with a better way to benchmark models in the open because it's largely uninteresting atm. And no, the solution isn't a benchmark that tests how far models get in a 2-day task. There's definitely a better way to host and/or manage benchmarks to actually gauge capabilities, at least for longer than a year.
18
11
192
22,119
RLM (recursive language model) "framework" is pretty cool, it's an LM (harness) that can programmatically manipulate its input as external data and recursively invoke itself (or other LMs) on selected portions of it the key difference from a "conventional" coding agent is that the entire input / intermediate results can live in an external Python REPL, while the root LM decides what computations to perform (mitigating the context rot) to make it less abstract, let's say we have the following user query: "among questions associated with users 123 and 456, how many should be classified as 'entity' questions?" (an entity question is something like "who is Albert Einstein?") let's say that the full dataset contains 5k entries: Date: ... || User: 789 || Instance: How do I bake bread? Date: ... || User: 456 || Instance: What is the capital of France? ... important bit here: the root LM does not receive these 5k entries in its prompt, the RLM harness stores them in a var called `context` steps in the RLM framework: 1. the root LM sees the user's query and knows that context exists (through system prompt), it then generates Python code: print(context[:2000]) the REPL executes the code and returns the first 2k chars (this partial observation is fed into the context) 2. the root LM now understands the dataset's format. next up it filters, e.g.: lines = context.splitlines() relevant = [ line for line in lines if "User: 123 ||" in line or "User: 456 ||" in line ] print(len(relevant)) let's say that returns 347, the root has reduced 5k entries to 347 (without polluting the context with irrelevant records) 3. next up the model can chunk up the relevant lines and recursively call itself: chunks = [ relevant[i:i+50] for i in range(0, len(relevant), 50) ] results = [] for chunk in chunks: result = llm_query( "Classify each question as entity or non-entity. " "Return the number of entity questions.\n" + "\n".join(chunk) ) results.append(int(result)) the harness executes seven sub LLM calls in this example, and the results are now in `results` var, e.g. [17, 21, 19, 23, 20, 18, 17] finally the model can return sum(results) as the result this is in stark contrast compared to the context-rot that would have happened without the RLM framework furthermore you can RL post-train the root LM inside this harness, they demonstrate much better generalization because at this level of abstraction many problems look the same - this is probably the most important bit work by @a1zhang and @lateinteraction !
24
9
172
7,728