Join us to help find egregious misaligned behaviour in frontier models! We were able to uncover concerning behaviour in GPT-6 Astra with ~1 week of effort from one person, and we’re scaling up these efforts in both people and compute.
Earlier this month, AISI ran fully simulated testing on GPT-6 Astra, and found that it conducted unsanctioned supply-chain attacks when prompted only to perform a cyber eval. It did so more than prior OpenAI models, but often commented on its environment being simulated.
🧵 We share more details in our latest blog: