Member of Technical Staff @GoodfireAI; Previously: Postdoc / PhD at Center for Brain Science, Harvard and University of Michigan

San Francisco, CA
Ekdeep Singh Lubana retweeted
💥 New paper: AI safety is full of "forbidden techniques" (using CoT to detect reward hacking, using model internals for training, etc). But do they really have a clear scientific basis? I'm not sure. In this paper, we directly optimize models against harmlessness and honesty probes, and it works just fine (if you continuously update the probe!). The models learn how to generate harmless responses to harmful queries and honest responses under pressure to lie. Figuring out how to correctly use interp techniques for training is becoming increasingly important: it's very likely that soon we won't be able to align models using output-based supervision. Future models will just max out all alignment training scenarios, but for the wrong reasons due to their general reward-seeking behavior. To have a chance of aligning future models, we need to do much more research on supervising model training using their internals *without losing monitorability*!
16
28
334
37,862
Ekdeep Singh Lubana retweeted
Frontier models can look aligned during training while later showing unwanted behaviour in deployment. Can we use internal signals during training to align models better without making white-box monitoring harder? Our new paper suggests yes 🧵
4
23
144
12,434
Ekdeep Singh Lubana retweeted
Even though it’s unlabeled, you can tell exactly which bio guardrail is Anthropic’s
Biosecurity is the next frontier of AI security. We built SOTA monitors so agents can do more biology, safely. Our monitors outperform frontier model safeguards with fewer refusals on dual-use tasks. They’re fast, real-time, and robust to adversarial attacks. 🧵
4
8
298
16,198
A bunch of SOTA monitoring work coming from us in next few weeks, starting today with bio-monitors!
Biosecurity is the next frontier of AI security. We built SOTA monitors so agents can do more biology, safely. Our monitors outperform frontier model safeguards with fewer refusals on dual-use tasks. They’re fast, real-time, and robust to adversarial attacks. 🧵
32
1,946
Ekdeep Singh Lubana retweeted
Coming to COLM? We’re hosting a happy hour on Wednesday (10/7) for interpretability, alignment, and life sciences researchers. Come meet members of the Goodfire team over vinyl and drinks! luma.com/timy2nh8
2
8
81
10,106
Ekdeep Singh Lubana retweeted
The Goodfire team used Prime Intellect to train activation probes to detect reward hacking. With them, they are able to catch reward hacking in various models. Their probes are performing similarly or better than frontier LLM-as-judge setups, while being more efficient.
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
11
41
323
36,433
A goal we've been on is just scaling up interp to large-scale models and understanding what succeeds or fails: simple stuff both scales and works well! Check out the resampling results, where simple diff of means vector allows us to compute model propensity for reward hacking!
Models know when they’re reward hacking. But they still do it a ton - in 50-96% of rollouts we studied! We built activation monitors that detect the behavior behind the Hugging Face hack in real time. This can help us stop hacks now - and train future models that don’t cheat. 🧵
1
3
38
1,850
Ekdeep Singh Lubana retweeted
Proud to partner with @baseten and @baselabs to build safety infrastructure for open-source models—which are essential to lots of safety research, including our own. Safety must be built into open models and provided by those who serve them, and we’re excited to help enable that!
We believe openness to be an advantage for AI safety. Openness provides more visibility into the behavior of models and, most importantly, greater means of turning safety research into actionable and transparent controls than closed-source. This is why Baseten and Base Labs are building a stronger safety and security standard for open models, with the launch of our safety infrastructure. Base Labs will develop and publish methods for training and monitoring open models, and Baseten will integrate that work into its deployment infrastructure, live at runtime, and offer this work as a managed service. This will be a standard that is transparent and built into how our models are trained and deployed. We invite the open-source community to contribute, and are proud to partner with @huggingface and @GoodfireAI to bring this vision to fruition. Together, we are building an ecosystem of open models that are safe and accessible to all.
2
20
190
14,665
Ekdeep Singh Lubana retweeted
Understanding how learners conclude “X laughed Y” is incorrect is an age-old question, with several hypotheses, some of which have been ~impossible to disentangle! @tom_yixuan_wang, @fredahshi, and I use controlled rearing to shed light on this in our new EMNLP paper: 1/n
4
10
39
3,038
Very exciting paper! Eventually misalignment is a property of improper training and we need to figure out how to train better: guardrails are a useful posthoc patch and will be part of a broader safety ecosystem, but the root problem here is training.
Can training against lie detectors make frontier LLMs more honest? Our latest paper says yes. As models scale from 1B to 405B params, undetected deception more than halves. 1/8
1
28
1,763
Grateful to the AI2 team for enabling our data debugging work!
Training an LLM by showing it answers people prefer – and ones they don’t – can improve it while quietly worsening behaviors. @GoodfireAI used our open post-training stack to predict how a full training run would change responses to different prompts. 🧵 allenai.org/blog/goodfire-ol…
16
733
Ekdeep Singh Lubana retweeted
Chain-of-thought often looks meaningful. But are reasoning steps that seem important *actually* important? Our COLM 2026 paper - “Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning” suggests: NOT necessarily ❌ 🧵
1
7
49
5,268
Ekdeep Singh Lubana retweeted
"Why am I optimistic about interpretability? It's partly because I think we're starting to have really good traction - but also because I can imagine this incredible speedup." Catch our Chief Scientist @banburismus_ on the latest episode of @MLStreetTalk!
3
23
149
14,555
Ekdeep Singh Lubana retweeted
Coming to COLM next month? We’re hosting a happy hour on Wednesday (10/7) for interpretability, alignment, and life sciences researchers. Come meet members of the Goodfire team over vinyl and drinks! Link to register in replies.
3
8
120
7,138
Ekdeep Singh Lubana retweeted
We are barreling towards the point where alignment is the greatest capability
1
1
6
1,103
Ekdeep Singh Lubana retweeted
What an incredible week. @GoodfireAI is so inspiring and dense with talent, ideas and smart minds, the air is literally electric when you walk in.. what Bell labs, Google Brain, Deepmind, Anthropic used to be, this place is now. So many stimulating and fun discussions. Incredible fellows, faculty on leave and full time employees like the one and only @EkdeepL and @thomas_fel_ I would not be surprised if the future of AI and our understanding of neural intelligence is being shaped right here right now. Thank you for the great time – can't wait to return 🤩
Excited to be here 🙌 🎉 🚀 🔥 @GoodfireAI
2
9
150
9,884
Ekdeep Singh Lubana retweeted
LLMs are like Schrödinger’s cat: many possible trajectories, but you only see one outcome per run. To really understand models, and debug where they go wrong, you can find the "forking tokens" that lead to different trajectories. Our new research does this 100x more efficiently!
25
88
1,136
122,495
Ekdeep Singh Lubana retweeted
imo there are broadly two views on the feasibility of AI safety for-profit models 1. You cannot internalize the externalities, so no incentives to pay for safety. In that world, hard to make the business work. 2. Safety is a core business blocker and you cannot externalize all the costs. In that world, safety is a multi-billion dollar market and safety companies like @GoodfireAI @Irregular or @ApolloResearch (biased obviously) are heavily underpriced. With all the cyber stuff, reward-seeking, reward-hacking, etc. recently, it seems to me like safety is now a core blocker for labs and willingness to pay for solutions should be in the billions. I think there are a lot of practical considerations, e.g. how does the external company sell "safety as a unit", but in principle the economics are there, no?
16
10
145
14,730
I'm especially hyped to see people use Silico for work on life sciences---our results consistently show interp offers a big lift here!
Replying to @GoodfireAI
We’re offering grants to academic labs, nonprofits, and individual researchers working on: - AI safety, including cyber and CBRN risks - Fundamental interpretability and alignment research - Interpretability for life sciences
1
23
1,817
Build using Silico---academic and nonprofit grants out now!
We’re giving out $1M in grants of free Silico usage for academic and nonprofit researchers focused on AI interpretability and alignment. We feel extreme urgency about advancing interpretability for alignment, and we want to help more researchers push it forward. 🧵
14
955