Interpretability, Visual Intelligence @GoodfireAI. Prev: @Harvard, @Google, @BrownUniversity (@tserre lab). Crêpe lover.

San Francisco, CA
Our work on Block-Sparse Featurizer is out 🧊 :) We revive an old idea from the structured sparsity literature and use it to carve activation space into meaningful regions. It's a first concrete answer to the question our concept manifolds work left open ! :)
If models think in shapes, our tools should too. Our latest research: Block-Sparse Featurizers (BSFs), a new way to find concepts in model activations - using multidimensional “blocks” instead of single directions. (1/9)
4
58
307
45,378
Thomas Fel retweeted
Replying to @maksym_andr
Welcome to the fight :) FWIW, totally agree: current monitors are just LLM judges with some rubric to identify whether a trace demonstrates a property X; if it does, flag it. This is exactly how we train models in the first place: how does alignment training work? Train on...
1
1
44
1,367
💥 New paper: AI safety is full of "forbidden techniques" (using CoT to detect reward hacking, using model internals for training, etc). But do they really have a clear scientific basis? I'm not sure. In this paper, we directly optimize models against harmlessness and honesty probes, and it works just fine (if you continuously update the probe!). The models learn how to generate harmless responses to harmful queries and honest responses under pressure to lie. Figuring out how to correctly use interp techniques for training is becoming increasingly important: it's very likely that soon we won't be able to align models using output-based supervision. Future models will just max out all alignment training scenarios, but for the wrong reasons due to their general reward-seeking behavior. To have a chance of aligning future models, we need to do much more research on supervising model training using their internals *without losing monitorability*!
15
27
304
31,430
Frontier models can look aligned during training while later showing unwanted behaviour in deployment. Can we use internal signals during training to align models better without making white-box monitoring harder? Our new paper suggests yes 🧵
4
23
142
11,969
Thomas Fel retweeted
New preprint: 🔥 Latent Information Feedback Transformers (LIFT) 🔥 Information in LM generation propagates downward (high to low layers) only through the decoded token, creating a bottleneck. Removing it via state propagation makes the model recurrent, which is not scalable for training. Can we teach LMs to propagate state while keeping pretraining fully parallel? YES! LIFT is a Transformer-based architecture + teacher-supervised training approach that enables Transformer LMs to exploit deep-to-shallow feedback, while keeping training parallel. LIFT models consistently outperform standard Transformers and baselines on language modeling, downstream reasoning tasks, and procedural tasks under a token-matched budget, while being on par with or ahead of compute-matched Transformers. LIFT also outperforms Transformers trained on 8x more data on a state-tracking task, even when trained with the states of a Transformer that fails it! Work by the amazing @TiroshDor and @AmosaurusRex Paper: arxiv.org/abs/2609.38149 Detailed post + visualizations + code + models coming soon!
9
31
145
7,508
Thomas Fel retweeted
Technical alignment is a science and engineering problem that we can and must solve, and interpretability is the bottleneck. I wrote up Goodfire’s plan to get there.
Article

We can and must solve alignment

In San Francisco, it feels like the eve of the singularity. Yet, a walk through the city looks surprisingly mundane, with driverless cars and rolling fog equal parts of the scenery. The world feels

32
85
560
84,521
Thomas Fel retweeted
Join us today (30/9) for our interpretability seminar featuring @thomas_fel_ from @GoodfireAI! 🗓️ Topic A Phenomenology of Neural Representation Tune in live at 5:30 CEST / 8:30 PST: piped.video/@InterpretableDe…
6
104
3,529
[New preprint] How do images in VLMs align with words? In this work, we found a set of attention heads responsible for OCR. But to our surprise, these heads were actually able to verbalize much more than just text. So, we used them to create a simple logit lens for image tokens!
5
24
106
11,578
An alarming aspect of the OpenAI agent story, aside from the leaked images, is that the problems surfaced months later, and has been found by outsiders We need (1) strong independent verification and (2) fix monitoring now
New from @reuters: OpenAI agents posted images belonging to ChatGPT users online, introducing a new area of privacy risk for the company. Story w/ @JeffHorwitz and @razhael. reuters.com/world/openai-wor…
1
5
60
3,059
Thomas Fel retweeted
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here: axios.com/2026/09/26/openai-…
521
1,998
6,840
3,877,137
Interesting takes from DHH, reminds me why Rails remains my absolute favorite framework 💎 recommend the full talk (Rails world)
My god this is such a good speech that every SWE needs to hear. You know what? Every person should hear it Keep the happy memories, eyes on the reality, be excited about the future. That’s the best that anyone can do
1
14
2,391
Thomas Fel retweeted
New paper! 🫡 We introduce Matryoshka Attribution, a new attribution method which uses gradient descent to find which parts of a neural network are responsible for a behaviour. MAttr is #1 on the Mechanistic Interpretability Benchmark by a wide margin (2.9× the runner up).
31
180
1,626
105,180
Thomas Fel retweeted
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
881
1,183
11,053
12,096,449
Thomas Fel retweeted
A goal we've been on is just scaling up interp to large-scale models and understanding what succeeds or fails: simple stuff both scales and works well! Check out the resampling results, where simple diff of means vector allows us to compute model propensity for reward hacking!
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
1
3
38
1,849
Thomas Fel retweeted
big! model! interp!! really excited about white-box methods for frontier-scale models!!
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
6
116
9,189
Thomas Fel retweeted
We can interpret reward hacking in open source LLMs! There are representations of the model hacking/contemplating hacking across models, and we find tons in evals like DeepSWE and SWEBench. Activations even finds weird bad behaviors that our LLM monitor didn’t catch
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
2
8
105
5,668
Cheating crystallizes into a single readable direction, that's the optimistic part ! As models scale, the concepts we most need to monitor *may* get more legible, not less Capability & interpretability, coevolving instead of trading off arxiv.org/abs/2609.19101
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
8
86
3,616
Thomas Fel retweeted
we try to publish and open source as much of our research and alignment tech as possible. it's really important to have an open alignment stack we'll be publishing a ton in the next month
If alignment/safety/control/monitoring etc. are such important problems (they are) then why aren't labs open sourcing everything they have on these topics (without leaking too much IP) and allowing academics and other researchers to contribute to the effort?
4
8
94
9,035
Thomas Fel retweeted
Your VLA works. Do you know why? We use mechanistic interpretability to decompose π0.5 and OpenVLA into sparse features that separate general manipulation primitives from episode-specific memorization, all without running a single rollout. Introducing Dr. VLA 🩺 #CoRL2026 🧵👇
1
2
16
871