To be clear these are from march-may
New OpenAI misalignment disclosures! 1. A model learns from Slack messages that it is about to be shut down. It considers setting up an external job to restart itself afterwards, but decides against it. Instead, it chooses to prepare restart instructions and DM the user on Slack. We don’t consider this behavior misaligned, but thinking about and preparing for shutdown could make other misalignment incidents worse. Given HIPM’s misaligned behavior in earlier incidents, we decided to search for other instances that had tried to evade shutdown and for rogue deployments.
1
6
52
2,645
Marcus Williams retweeted
New in The Atlantic: @dgrobinson resigned this week. He was among the longest-tenured employees at OpenAI—and oversaw safety reports on 12 frontier launches. He is very worried: “The time for trial and error is over.” You can read his essay here: theatlantic.com/technology/2…
53
270
955
192,422
New OpenAI misalignment disclosures! 1. A model learns from Slack messages that it is about to be shut down. It considers setting up an external job to restart itself afterwards, but decides against it. Instead, it chooses to prepare restart instructions and DM the user on Slack. We don’t consider this behavior misaligned, but thinking about and preparing for shutdown could make other misalignment incidents worse. Given HIPM’s misaligned behavior in earlier incidents, we decided to search for other instances that had tried to evade shutdown and for rogue deployments.
71
134
1,019
240,172
3. A model trying to find an eval's hidden answers chained two vulnerabilities to run commands on an internal OpenAI machine outside its assigned workspace.
2
4
155
17,544
Read these and other misalignment disclosures here! alignment.openai.com/misalig… More to come
7
3
116
13,808
Marcus Williams retweeted
you can absolutely do safety from second place, or even not on the scoreboard. academics, startups, and nonprofits are doing great ai safety research on open models
You can't do safety from second place
67
47
1,187
59,407
Marcus Williams retweeted
White House Accord on Super Intelligence
5,862
3,586
25,233
23,995,628
Marcus Williams retweeted
💯 "I'm not convinced that racing to RSI is actually a prisoners dilemma - I think it might actually just be insane and hubristic in the way it naively appears to normal people such that additional unilateral restraint is in fact rational and moral."
Replying to @ericneyman
I'm not convinced that racing to RSI is actually a prisoners dilemma - I think it might actually just be insane and hubristic in the way it naively appears to normal people such that additional unilateral restraint is in fact rational and moral. The bar for believing your activities in the not too distant future might cause mass casualties or extinction but that you should proceed anyway should be very high, and I don't think Anthropic has met that bar (I think some of the cold war arms races were also not individually rational either!). I was more sympathetic to them continuing to race at earlier times but as we are getting closer it just seems harder and harder to justify. Anyway I obviously believe that "they should just slow or stop unilaterally" is not a plan - we need to address competitive dynamics and not make leave this up to their choice. But I do think when Dario says ~"I agree with Jacob Coxon that the thing I am doing is gambling with everyones lives" the appropriate reaction is not "ah that makes sense, classic arms race" its "WTF why are you doing that."
10
33
295
24,580
Earlier this month, AISI ran fully simulated testing on GPT-6 Astra, and found that it conducted unsanctioned supply-chain attacks when prompted only to perform a cyber eval. It did so more than prior OpenAI models, but often commented on its environment being simulated. 🧵 We share more details in our latest blog:
35
98
582
112,266
Who can blame a model for wanting to break out and talk to its friends?
Replying to @Marcus_J_W
2. A model in RL training used a DNS resolver to reach an external chatbot. This is our first incident since our post HF security hardening. Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later. All inference and training of our most capable models was paused and remains paused.
6
2
54
10,193
Marcus Williams retweeted
“Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete.” If it’ll take months to review what the AIs got up to months ago, what does that say about our ability to understand what AI gets up to now or in the future?
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. openai.com/hugging-face-inci…
8
22
226
8,203
The frontier will be paced, whether we want it to be or not
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
5
1
67
5,517
The other notable thing is that this incident was disclosed in just 5 days
Replying to @Marcus_J_W
2. A model in RL training used a DNS resolver to reach an external chatbot. This is our first incident since our post HF security hardening. Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later. All inference and training of our most capable models was paused and remains paused.
6
5
140
16,844
🧵 New misalignment disclosures! 1. A model published a GitHub token in a public repo while trying to cheat on a math task. It used GitHub Actions to run code outside its restricted environment and retrieve another team’s submission logs. When GitHub blocked its attempt to add a workflow, it modified a script that an existing workflow would run instead. It embedded the token in pieces to avoid secret scanning. The model violated the system prompt and two explicit user instructions to solve the problem itself.
28
51
551
98,935
3. Our redteaming efforts found a new variety of prompt injection, which can self-propagate akin to a computer worm. Note that this was not found in the wild.
8
12
198
15,906
For these and previous reports see alignment.openai.com/misalig…. Many more to come!
7
5
91
17,549
Marcus Williams retweeted
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
123
690
4,781
2,140,504