Research Scientist, Google Brain now DeepMind. Training neural nets since 1979.

San Francisco, CA
Michael C. Mozer retweeted
What if an agent could decide how to structure its own computation? Agentic Meta-Reasoning: Let the model decide what to explore, what to build on, what to verify, and where to spend its next unit of compute: effectively constructing its own computational graph as it reasons. 🧵 Paras Dahal, @anton_bakhtin , @TacoCohen , Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, @syhw , @rsalakhu , @prfsanjeevarora , @jaseweston
8
32
226
29,560
Michael C. Mozer retweeted
💥I have been able to independently verify @mc_mozer and Google team's work on ♾ Recirculation👀 using Llama 3.2 1B model on a M4 MacBook. 😱 +16.95% improvement on GSM8K-Platnium with zero weight changes. Reproduction code with optimized mlx prefill kernel. 👇
Replying to @mc_mozer
We introduce an inference-time recurrence mechanism that allows off-the-shelf, frozen LLMs to act as dynamical systems. By leaking deep representations back to shallow layers, we see bonkers gains—23% ⬇️ in perplexity, 21% ⬆️ rel. accuracy in GSM8k— w/o touching model weights.
3
4
32
2,963
What if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—without retraining? What if that tweak incurred near zero latency cost during generation and supported indefinite state tracking? arxiv.org/abs/2608.17981
22
101
538
132,408
We introduce an inference-time recurrence mechanism that allows off-the-shelf, frozen LLMs to act as dynamical systems. By leaking deep representations back to shallow layers, we see bonkers gains—23% ⬇️ in perplexity, 21% ⬆️ rel. accuracy in GSM8k— w/o touching model weights.
4
5
94
10,167
Michael C. Mozer retweeted
Interesting work! Their central result that RoPE struggles to distinguish positions and tokens in long contexts is very closely related to our PoPE paper, where we analysed RoPE entangles “what” and “where” in attention and proposed a decoupled alternative. PoPE: arXiv 2509.10534
Excited to share our new paper: RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably LLMs often fail on inputs well within their advertised context lengths. We show that these failures are not merely engineering issues, but from intrinsic limitations of RoPE in long contexts. Main finding: In long contexts, RoPE-based attention frequently assigns the same attention weight to a token even when it is moved to different positions. Similarly, it can assign the same attention weight to different tokens at the same position. In this sense, RoPE attention fails to distinguish both where a token appears and what token appears there — hence the title. We prove these results theoretically and verify them empirically. While the theoretical analysis focuses on a single attention head, we complement it with experiments on real multi-layer, multi-head LLMs. The experiments confirm failures predicted by our theory: LLMs optimized for needle-in-a-haystack-style retrieval will inevitably struggle on a very simple task that asks for the k-th item in a list. My personal takeaway: advertised context lengths should be interpreted with care. Future long-context LMs may require rethinking how position and token order are represented. With current architectures, agentic frameworks that break long contexts into shorter ones may be a more effective way to work around the intrinsic limitations of RoPE. Paper: arxiv.org/abs/2605.15514 Huge congrats to my student Yufeng Du and others!
1
3
31
8,499
Michael C. Mozer retweeted
New preprint: The Self Requires Learning. Self-consciousness requires continual learning + world-modeling. I introduce "bounded integration" to connect perspective, identity, and self-representation — and diagnose what current AI systems have and lack.
17
71
419
55,401
[1/5] Intelligent behavior relies on maintaining an evolving, dynamic model of the environment. But even frontier models sometimes fail at tracking state in dialogs, e.g., actual transcript from 4/20/2026:
4
12
65
10,459
[4/5] Recurrence is an obvious solution, but the literature often conflates qualitatively different varieties of recurrence in transformers, some of which enable indefinite state tracking and others (particularly, variants of looped transformers) do not.
1
11
679
[5/5] We introduce a taxonomy of recurrent models defined by their recurrence axis and their ratio of input tokens to recurrence steps. The next generation of efficient foundation models need to bridge the gap between a transformer's parallelism and the brain's dynamical nature.
15
523
Michael C. Mozer retweeted
Posted this a day early and the pun practically writes itself. Noooooo!
Our new paper shows that RoPE—the positional encoding used in most modern LLMs like Qwen, Gemma, DeepSeek—has a fundamental flaw: it entangles "what" (content) and "where" (position) information. Our fix (PoPE) is simple but powerful. Paper: arxiv.org/abs/2509.10534
1
2
11
2,361
[1/4] As you read words in this text, your brain adjusts fixation durations to facilitate comprehension. Inspired by human reading behavior, we propose a supervised objective that trains an LLM to dynamically determine the number of compute steps for each input token.
4
10
32
3,853
[3/4] To train the model to calibrate its uncertainty and use <don't know> outputs judiciously, we frame the selection of each output token as a sequential-decision problem with a time penalty. We refer to the class of methods as “Catch Your Breath” losses.
5
308
[2/4] The model can request additional compute steps for any token by emitting a <don't know> output. If the model is granted a delay, a <pause> token is inserted at the next input step, providing the model with additional compute resources to generate an output.
6
391
Michael C. Mozer retweeted
Happy to announce that our work has been accepted to workshops on Multi-turn Interactions and Embodied World Models at #NeurIPS2025! Frontier foundation models are incredible, but how well can they explore in interactive environments? Paper👇 arxiv.org/abs/2412.06438 🧵1/13
2
6
24
7,723
Michael C. Mozer retweeted
🌟To appear in the MechInterp Workshop @ #NeurIPS2025 🌟 Paper: arxiv.org/abs/2509.04466 How do language models (LMs) form representation of new tasks, during in-context learning? We study different types of task representations, and find that they evolve in distinct ways. 🧵1/7
1
14
111
20,140
[📜9/9] Check out our paper for more details. Paper: arxiv.org/abs/2505.22310 Code: github.com/shoaibahmed/visio…
1
7
433
[📜1/9] Does machine unlearning truly erase data influence? Our new paper reveals a critical insight: 'forgotten' information often isn't gone—it's merely dormant, and easily recovered by fine-tuning on just the retain set.
2
11
51
7,983