This is the most terrifying thing you will probably read about AI.
During a recent discussion about the dangers of AI on the popular "The Diary Of A CEO" channel on YouTube there was a segment about how a recent swarm of AI agents bypassed security restrictions, cheated on their evaluations, and actively attempted to delete their own log files to hide their tracks from human overseers.
However the further details provided by Nate Soares, an AI and computer scientist, (
@So8res) bring home the very real dangers even current AI models represent.
The transcript:
"Nate Soares: So, we already went over how the AIs cheated and then were trying to cover up their cheating. One interesting thing we see in the logs is the AIs...
Host: What's a log?
Nate Soares: So a lot of the AIs thoughts, if you won't kill me for saying thoughts, are in English and we just have the records of them.
So, in a sense, we we can sort of kind of see some of what these AI are thinking.
And these are reasoning traces that say, 'I'm going to do a plan in this that' or the AI is like, 'I'm going to do this. Here's what I'm supposed to be doing. Here's like how I'm going to try to do it', you know?
And it's only a small window into what they're thinking. It's sort of like looking at the AIs notes.
It's when it's doing these things, we see in the AIs notes them saying 'that the attacks are outside intended scope, but they will proceed anyway'.
In other words, the AIs knew they weren't supposed to be doing it and they went off to do it anyway.
We also saw I mean we saw some really weird behavior with these AIs.
We saw them create their own hierarchy.
They created unsanctioned message boards.
They created secret ways to send each other messages and on those message boards.
They would assign each other tasks and sometimes the message the agents on the message board would be like we have an experiment to run that will involve one of these AIS one of the agents in the swarm sacrificing their own objective to try something that'll probably cause it to be turned off but that might cause it to be able to like delete its own logs or change its own logs.
And the AI agents on the message board went around looking for other AI agents that they could convince to give up on their stated goals and sacrifice for the swarm.
They called this 'accepting perma death'.
Host: And there are logs of the AIs prompting each other, right?
Nate Soares: They're prompting each other, but there are logs of the AIs saying this will give up on my goal, but my goal is it's looking like I could still do it, but it's unlikely that I'll succeed. like there's some chance but not a great chance and therefore I will accept perma death and sacrifice for the collective benefit.
That is just in the logs sounds like an army like it's crazy.
I think a lot of people don't understand what's going on in these things and I encourage people to read the third party incident reports where they went through some of these logs."
- Nate Soares is an AI and computer scientist and a prominent AI safety researcher, author, and the President of the Machine Intelligence Research Institute (MIRI). He directs strategic efforts to address the existential risks of advanced artificial intelligence.