Using interpretability to understand, learn from, and design AI.

San Francisco
Pinned Tweet
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
31
100
849
153,222
Goodfire retweeted
Reward hacking, and the race to see inside AI models: my conversation with @eric_ho, CEO of @GoodfireAI Is this the beginning of the interp exponential? 00:00 Cold open & intro 01:06 "Amoral students with an absent teacher" 02:49 What is reward hacking? 04:13 Models cheat up to 96% of the time 06:08 Do models know they're cheating? 08:10 The Hugging Face hack: why tests missed it 11:35 "The dumbest models we'll ever deal with" 11:58 Should AI slow down? 13:56 Mechanistic interpretability 101 16:06 Models are grown, not built 19:10 How labs do alignment today 21:30 Is this an alignment crisis? 23:26 Why agents changed everything 25:38 Chain-of-thought monitoring is fading 27:49 Neuralese: AI that stops thinking in English 29:28 A model caught evading its own monitor 30:42 Open vs. closed models 32:11 Activation monitoring in production 35:00 Learning from superhuman AI 36:15 The most underrated field in AI 38:24 What is a probe? 39:52 What is steering? Golden Gate Claude 40:53 Why build Goodfire outside the labs 42:46 Do you need frontier model access? 44:16 Inside the paper: an MRI for the model 47:28 Dialing sycophancy up and down 49:16 Probes vs. chain-of-thought monitors 51:44 Cutting monitoring costs by 90% 53:28 So what can we do about it? 55:41 Giving gradient descent a choice 57:24 RL from feature rewards 1:00:03 Silico and Goodfire's business 1:01:25 A new Alzheimer's biomarker 1:03:38 Decoding neural networks by 2028? 1:06:00 What engineers can do tomorrow 1:07:35 Are we at risk? "I want people to believe" Great group of folks and friends involved @deedydas @MenloVentures @fendien @Work_Bench @BCapitalGroup @AnthropicAI
8
5
19
8,257
Biosecurity is the next frontier of AI security. We built SOTA monitors so agents can do more biology, safely. Our monitors outperform frontier model safeguards with fewer refusals on dual-use tasks. They’re fast, real-time, and robust to adversarial attacks. 🧵
6
23
137
32,267
Our method is also 3-5x more robust to adversarial attacks than well-established screening methods. And this improves as biology models become more capable. The same advances in biological understanding that unlock new research can also be what keeps it safe!
1
15
917
Our CEO @eric_ho just published Goodfire’s view on why we can and must solve alignment - and our plan to get there.
4
9
158
9,176
Goodfire retweeted
Technical alignment is a science and engineering problem that we can and must solve, and interpretability is the bottleneck. I wrote up Goodfire’s plan to get there.
Article

We can and must solve alignment

In San Francisco, it feels like the eve of the singularity. Yet, a walk through the city looks surprisingly mundane, with driverless cars and rolling fog equal parts of the scenery. The world feels

33
85
561
84,954
Are SAEs dead? Will they save us from neuralese? Should I just use a probe? We get these questions all the time. Part 2 of our educational series on applied interpretability explains what SAEs are good for, when *not* to use them, and what to use instead. 🧵
11
52
439
36,367
Training a good SAE is harder than it looks: choosing layers, collecting activations at scale, tuning width and sparsity, labeling millions of features… Our blog post walks through these choices. Or just describe your goal to Silico, our interp agent, and it handles the rest.
1
28
1,782
Goodfire retweeted
opus music videos are so good that you can learn about @GoodfireAI mech interp research and get llm psychosis at the same time
Replying to @GoodfireAI
We’re optimistic that every training run can be monitored and that the current rates of reward hacking may soon be a thing of the past. Read the full post + paper: goodfire.com/research/reward…
9
27
317
26,534
Coming to COLM? We’re hosting a happy hour on Wednesday (10/7) for interpretability, alignment, and life sciences researchers. Come meet members of the Goodfire team over vinyl and drinks! luma.com/timy2nh8
2
8
81
10,115
The Goodfire team used Prime Intellect to train activation probes to detect reward hacking. With them, they are able to catch reward hacking in various models. Their probes are performing similarly or better than frontier LLM-as-judge setups, while being more efficient.
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
11
41
324
36,448