To develop frontier AI safely, developers must be able to show that their models are not scheming, i.e. covertly working against them in pursuit of unintended goals.
New post: four claims any scheming safety case must make, and the resources embedded evaluators need to verify them 🧵