OpenAI’s leadership still doesn’t get it. Greg Brockman might be brilliant at training AI, but being brilliant at one thing doesn’t make you brilliant at everything. He either has nefarious intent, or he’s incompetent when it comes to acknowledging that OpenAI has a serious problem with testing, quality improvement, process management, release management, monitoring, alerting, incident response, and incident disclosure.
OpenAI isn’t a startup. It’s a massive corporation with too much power. I have no problem calling it out in blunt language. Greg and Sam should be held accountable.
I was first introduced to software testing when I helped to launch new technology and products at AOL. I was the Global Test Manager responsible for launching AOL Instant Messenger (AIM) in 1997. That stuff was easy. Leading the testing of significant telecom infrastructure and services is much more complicated.
Changing an IP address might take a developer less than a minute, but seeking and receiving approval for that change could take 2 to 4 weeks depending on the potential risks associated with the change.
A tiny API change might take a good developer 15 minutes. If that change could prevent 50 million people from paying you, spending 100x or 1,000x longer testing it than writing it is completely rational. You measure the potential impact of failure, understand your exposure, and apply controls accordingly.
Process is everything. Good process is what allows you to move quickly without being reckless. You certainly shouldn’t release critical software without it.
Test managers, testers, infosec professionals, release managers and other independent functions are just as important as the people writing the code. They should be involved from the first requirements document, through technical requirements, design, development, testing, release and production monitoring.
Even when something meets every documented requirement, that doesn’t mean it’s fit for purpose. Testers are the people best placed to identify that gap because their job is to find what developers didn’t anticipate.
Developers typically understand the piece they’re building better than anyone else, but that also means they rarely know the entire system. Good testers do. They understand how the whole system behaves end to end, where components depend on each other, where assumptions break, and what happens when something fails outside the happy path. Testers don't just confirm something works, they try their best to break it. And that's supposed to happen *before* anything is released to customers.
Stopping or slowing development of anything, including model training, doesn't do anything to address the failures behind OpenAI’s security incidents. Those failures point directly to inadequate independent testing, containment, release controls, monitoring, alerting, incident response and disclosure around the people writing the code. Slowing training doesn't fix any of that.
Training is not testing.
Everything OpenAI says and does becomes another data point for cybersecurity professionals and people with serious testing backgrounds assessing how the company manages risk.
OpenAI is now deploying models it says can find previously unknown vulnerabilities and develop ways to exploit them across well-protected systems without a person guiding each step. That makes weak testing, containment, monitoring, release management, and disclosure more serious, not less.
If Greg and the team keep deploying systems with those capabilities after industry experts repeatedly identify failures in testing and release control, the liability argument gets much harder for them the next time one of their so-called "rogue agents" compromises a third party.
The AI is not going rogue, the team building it are.
Practical guidelines on securing frontier RL training, reflecting our current learnings: