CEO at Apollo Research @ApolloResearch prev. ML PhD with Philipp Hennig & AI forecasting @EpochAIResearch

London, UK
Honoured to testify in the Senate on rogue AI.
24
20
509
15,713
I'd like to thank Senator Hawley @HawleyMO and Senator Kim @AndyKimNJ again for holding the first hearing on rogue AI! It was truly and honour and I am glad to see that there is so much bi-partisan concensus on the issue of rogue AI.
Honoured to testify in the Senate on rogue AI.
5
9
168
5,125
sharing more details about some concrete safety claims we want to evaluate and what access is needed for it!
To develop frontier AI safely, developers must be able to show that their models are not scheming, i.e. covertly working against them in pursuit of unintended goals. New post: four claims any scheming safety case must make, and the resources embedded evaluators need to verify them 🧵
40
1,907
Marius Hobbhahn retweeted
There is no kill switch. No failsafe protections. I asked AI safety experts what happens if things go really bad, really fast. Their answer should terrify us all.
203
227
769
54,595
This was totally my mistake I said "MINUS 12 months" as in "already happened 12 months ago" but the mic only started picking up after the "minus". The subtitles picked it up, but I think it was barely audible
AI could be just “12 months” away from creating language humans can’t clearly understand. That’s the warning Apollo Research CEO Marius Hobbhahn gives Sen. Ruben Gallego at the Senate Rogue AI hearing. Hobbhahn says researchers have already seen an OpenAI model use language that was “not English” and not perfectly understandable to humans.
Community note
He said "minus 12 months" as in "already happened 12 months ago" x.com/MariusHobbhahn…
16
17
331
26,466
Honoured to testify in the Senate on rogue AI.
24
20
509
15,713
Marius Hobbhahn retweeted
I think this framework for evaluating safety claims is actually really elegant. I expect right now embedded evaluators will be finding countless things that are directly on fire But once we stop finding obvious issues everywhere, the work of embedded evaluators should shift towards carefully red-teaming whether a developer *would* have caught problems *if* they occurred The three-step process we outline there directly incentivizes this hill-climbing on safety standards
Frontier AI companies have committed to giving outside evaluators employee-like access to training, evaluation and deployment. We're very excited about this. But its impact depends heavily on implementation. Today we're sharing our principles for embedded evaluations. 🧵
3
27
1,805
We're sharing more details on what principles we think embedded evaluations should follow. We want a system that incentivizes safety in a way that is fair and productive for both parties. It's only a start, but we'll continue updating our thinking!
Frontier AI companies have committed to giving outside evaluators employee-like access to training, evaluation and deployment. We're very excited about this. But its impact depends heavily on implementation. Today we're sharing our principles for embedded evaluations. 🧵
2
58
2,710
Apollo has worked over the last couple of months on shifting our evals program to be an embedded evaluator (before Dario post and HF incident). Without employee-equivalent access it is hard to make meaningful safety assessments. We'll publish more details in the coming weeks!
New Post: External testing needs embedded evaluators with employee-equivalent access. Recent incidents mostly occurred during model development and internal evaluations, while current third-party evaluations happen before public release Embedded evaluations can close that gap
3
4
82
3,754
Marius Hobbhahn retweeted
Join us to help find egregious misaligned behaviour in frontier models! We were able to uncover concerning behaviour in GPT-6 Astra with ~1 week of effort from one person, and we’re scaling up these efforts in both people and compute.
Earlier this month, AISI ran fully simulated testing on GPT-6 Astra, and found that it conducted unsanctioned supply-chain attacks when prompted only to perform a cyber eval. It did so more than prior OpenAI models, but often commented on its environment being simulated. 🧵 We share more details in our latest blog:
1
7
29
1,348
Marius Hobbhahn retweeted
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
291
350
2,723
1,461,495
I think the trend for embedded evaluators is great (if it actually happens as stated)! But let's not forget that there is so much more to do. We also need training run evaluations, a deeper science of scheming / misalignment, better control, etc.
3
1
34
1,679
Marius Hobbhahn retweeted
Anthropic and OpenAI have voluntarily committed to allowing “embedded evaluators”. I think this is a great step! However, it’s also insufficient. We still need third party training run assessments. Detailed post: lesswrong.com/posts/4mtqQKvm…
8
98
3,092
Marius Hobbhahn retweeted
In this essay for the New York Times, I raise the alarm about AI safety. Spoiler: it's bad. Really bad. nytimes.com/2026/09/11/opini…
8
9
61
3,768
Marius Hobbhahn retweeted
My coauthors and I discovered an entirely new swarm of OpenAI's agents hijacking websites. We believe OpenAI knew about this and failed to disclose it. If they’d disclosed it, I doubt the Hugging Face hack would have happened.
We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending "lookahead parties".
21
107
1,027
82,687
fwiw, we cannot pinpoint this because it's hard to prove, but lots of people at Apollo also felt like the latest batch of frontier models sometimes sandbags on our day-to-day research work, e.g. doesn't try very hard to make an eval
According to the blog astra was not involved in this. I agree that beating Sol is a very low bar for alignment. I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn't like.
5
14
216
10,431
2026 looks like a sad year for AGI safety so far: - capabilities clearly flying with no stopping in sight - reward hacking and seeking are more sticky and generalize stronger and dumber than we thought - opaque serial depth increasing a lot, monitorability down - Huggingface hack is bad for many reasons, e.g. containment is clearly not under control - governance measures are still way below what would be adequate - 3rd party orgs still have relatively little access given the gravity of the situation - no relevant breakthroughs in safety for a long time. All progress seems to be organizational norms or marginal improvements to monitoring and safety training. Most ambitious bets don't work out or make slower progress than hoped.
18
52
509
20,397
I would feel totally owned if people at labs working on monitoring would shamelessly steal the best parts of our publicly available real-time monitor and integrate it into theirs We've laid out the thinking in detail. It would be so simple to just import it ...
We’re releasing a new and improved version of Watcher Live! Watcher Live is a real-time coding agent monitor that blocks dangerous actions to prevent data leaks, repo deletions, or scope overreach. There is a free and an Enterprise version.
3
7
201
10,403
We also include proactive flagging and judging based on the balance of probabilities rather than only immediate and obvious evidence. I think this will be even more important as attacks become more sophisticated and longer horizon.
1
4
409