Our goal is to secure frontier AI systems from development, to deployment and governance.

London
We're now @ApolloResearch. Same team, same mission. A better handle that reflects that research is at the core of everything we do. If you've been following us as @ApolloEvals, nothing has changed.
1
49
8,519
To develop frontier AI safely, developers must be able to show that their models are not scheming, i.e. covertly working against them in pursuit of unintended goals. New post: four claims any scheming safety case must make, and the resources embedded evaluators need to verify them 🧵
4
8
66
5,366
Access limited to a scope agreed in advance can't account for unforeseen risks, like a new training method with a novel failure mode. Sensitive data can stay on developer infrastructure with blocked egress, and algorithmic details can be protected by strict NDAs.
1
5
271
We're excited about embedded evaluations. Any developer should be able to show its models are not scheming, and we see these four claims as a starting point for what that takes. Full post: apolloresearch.ai/blog/towar…
1
7
264
Reducing risk means catching problems early. Warning signs like models subverting their training or evading monitoring could appear long before release. So evaluators need ongoing access to training data, rollouts, checkpoints and internal deployment.
1
7
230
Public reports come out on a fixed schedule. They list every claim and verdict, as well as the access the evaluator asked for and received. If the developer disagrees, it can add a statement, but it cannot veto the verdict.
1
5
303
We see this as a starting point and expect the methodology to evolve. We welcome feedback from developers, other evaluators and policymakers, and we hope the first embedded evaluation agreements set a high bar for those that follow. Full post: apolloresearch.ai/blog/princ…
8
292