In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power. In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
248
576
3,289
1,825,646
Daniel Kokotajlo retweeted
Sarah is Head of Public Policy at Anthropic. My read: she picked up on Anthropic’s implicit worldview, but didn’t realize she wasn’t meant to say the quiet part out loud, since EAs at Anthropic are still claiming to care about the kind of safety you can do from second place.
You can't do safety from second place
32
35
921
79,086
Daniel Kokotajlo retweeted
This is terrible and a clear effort to undermine safety research, whistleblowing, and the raising concerns policy.
Reporting from the WSJ. That fact that three alignment/safety researchers left this morning was already being discussed here on X.
3
7
138
12,818
Daniel Kokotajlo retweeted
As former comms staff at a frontier lab, I'll point out that only 5 of 12 interviews on frominside.ai/ are from current employees. Meanwhile 1,386 current frontier lab employees (~7%) signed the Pacing the Frontier letter. Even that number likely understates safety concern within the labs. Everything you're seeing in the public discourse should be understood as muted relative to the internal reality. I can say this with confidence: I was on the team responsible for making sure of this. I'm used to being behind the scenes and I've never done an interview before, but it felt important to bring more honesty to the conversation.
Replying to @PalisadeAI
@v_maini was the co-lead of AGI Deployment strategy at DeepMind. "The gap between Neanderthals and Homo sapiens in intelligence is far narrower than the gap in intelligence we should expect between humans and advanced AI systems to be." piped.video/watch?v=1S_Ay87d…
6
23
189
21,282
Daniel Kokotajlo retweeted
Replying to @ben_j_todd
AFAICT "doomer" gets used to refer to everyone who thinks there's a real risk, so 63% of Americans are "doomers".
2
1
66
2,828
Daniel Kokotajlo retweeted
Some people will look at misalignment incidents and insist that these are akin to bugs in traditional software. This is an actively bad analogy, because playing whack-a-mole with examples of misalignment (as one might with software bugs) not only fails to resolve the underlying problem but may in fact make it *worse* by making it harder to detect or even, depending on how you do the whack-a-mole, teach the machine to deliberately hide misalignment. This is not how traditional software works, and those who insist “it’s just like fixing bugs in software” are confidently applying a lossy analogy that confuses more than it clarifies.
I am more optimistic than this. I think it’s primarily a science problem, and that it’s possible to make progress in that science. But no, while there is engineering we can do today to improve AI safety, it is not fundamentally an engineering problem.
40
57
500
33,452
Daniel Kokotajlo retweeted
Yup, OpenAI delaying public deployment while continuing internal development and deployment increases the size of the internal/public gap, which hampers appropriate societal responses. aifuturesnotes.substack.com/…
Replying to @ShakeelHashim
If you're most worried about risks from Astra 6.1, it's good. If you're much more worried about future models, it's bad: risky external deployments have small downsides and they increase visibility into capabilities & propensities. Internal but not external is the worst scenario.
4
12
107
12,911
Daniel Kokotajlo retweeted
Dude, so you’re telling us that multiple AI companies lost control of their models, hacked other companies and government agencies—exposing the AI companies and employees to potential criminal prosecution—as a regulatory capture scheme? DO YOU KNOW HOW LUDICROUS YOU SOUND?
These stories are a fascinating look into how major corporations are leveraging fear and a naive (or complicit) media to peddle hysteria through carefully crafted stories where they face zero pushback or journalistic curiosity. Let's break this story down. Anonymous "sources" told Axios that OpenAI and Anthropic are investigating TENS OF THOUSANDS of incidents of "problematic" actions taken by their models. Notice who is talking… it’s not a whistleblower. Seven times the piece says "sources said" or "sources told Axios." If this were about transparency and safety, these companies should own it — put the count, the definitions, and the severity table in a blog post under their own names. But if it’s not that, and this is instead an unsanctioned leak of information—if someone inside believes these corporations were being reckless, then why is this some routine “scoop” and not some investigative piece about whistleblowers? The big tell here is that there’s no explanation why the reporter is giving the source anonymity. There’s no “out of fear of reprisal” or “who insisted on anonymity to discuss internal deliberations.” Instead, all we get is a story describing the corporations' own internal investigations, in the their own vocabulary, while Anthropic's figures come straight from its own public system card. OpenAI's own disclosure on Friday: "dozens" of organizations notified, and the "vast majority" of what it reviewed were mundane research tasks. Anthropic's own system card: 1.5%. Coincidentally, this comes just DAYS after both corporations released new models—Claude Opus 5.5 and GPT-6 Sol and Luna—on the same day. Those products are on the market right now. This is like a car manufacturer adding a new self-driving car to the market while the company is investigating TENS OF THOUSANDS of incidents where the self-driving technology failed to recognize stop signs and red lights, leading to the car unsafely (and illegally) barreling through intersections. If that’s actually the story being broken here, this is a Pulitzer Prize-worthy piece of investigative journalism! …But more likely, it’s just a corporate comms set up story, with the goal of leveraging hysteria and political backlash to protect their market position? Why weren't these questions asked? ❓How many of the tens of thousands are designed failure modes — tests built to make the model misbehave? ❓ How many are production agents vs. unreleased checkpoints vs. research swarms? ❓ What is the severity distribution — log scrape vs. credential theft vs. live compromise? ❓ What was the count, and the rate, on the previous generation? ❓ Why is Anthropic's own figure — escape attempts in 1.5% of runs on a sandbox test built to provoke them, all low-severity and self-reported, ~85% fewer than Opus 5 / Mythos 5.1 — being sold as evidence of lost control? Anthropic specifically and Dario spent September selling a "pause," "independent evaluators," and a One World Order, while unironically announcing a new wet lab in SF and shipping new product. If the number means what the headline says, stop shipping and open the logs. If it doesn't, stop laundering political objectives through anonymous sources and calling it a scoop. The media should smarten up and grow some sincere curiosity.
260
202
1,157
117,125
Daniel Kokotajlo retweeted
In 'well when you put it like that' news, here's the Florida Attorney general asking for a preliminary injunction to stop OpenAI from doing more AI R&D.
23
160
1,531
128,864
Daniel Kokotajlo retweeted
Could automating AI R&D radically accelerate AI progress in an “intelligence explosion”? Preliminary evidence suggests that it could. In a new paper with authors across academia, civil society, and frontier AI companies (including @dawnsongtweets @merettm @jackclarkSF @Yoshua_Bengio @geoffreyhinton ), we assess the evidence & offer policy recommendations 🧵
27
101
432
107,192
Daniel Kokotajlo retweeted
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there.
Article

Its not just the f*cking sandbox

A lot of the perspective on all the AI incidents has been shared from the outside in, and little has been said from the inside looking out, through the lens of a security person living through it.

444
412
2,896
1,373,325
Daniel Kokotajlo retweeted
Why didn't more AI safety people advocate for AI pause/stop earlier? Well, aside from sociological reasons, in order to do that, you had to think that none of the following would work out (be feasible, safe, have low enough safety tax) in time, which takes a degree of skepticism almost no one could muster. (Or object to AGI/ASI on non-consequentialist grounds, but there was apparently a very high correlation between consequentialism and early interest in AI safety.) Friendly AI (CFAI) Coherent Extrapolated Volition (CEV) Metaphilosophical AI Tool AGI Value Learning Oracle AI Agent Foundations Corrigible AI Quantilizers Approval-Directed Agents Human Imitation Cooperative Inverse RL (CIRL) Iterated Distillation and Amplification (IDA) Task-Directed AGI RL from Human Feedback (RLHF) Impact Regularization Mechanistic Interpretability AI Safety via Debate Recursive Reward Modeling Comprehensive AI Services (CAIS) Infra-Bayesianism Alignment by Default Natural Abstraction Eliciting Latent Knowledge (ELK) Shard Theory Constitutional AI AI Control Weak-to-Strong Generalization To recenter the sociology, it was much easier to build a career out of being bullish one or more of these approaches, than out of general skepticism. (I've been independent, financially and otherwise, throughout my participation in AI safety, which was perhaps not a coincidence from being the only AI pause/stop advocate for a long time.)
This actually left out the most unique thing I did. (Each item on Andreas's list was done by at least one other person besides myself.) Starting in 2004, I concluded that Friendly AI would be too hard/risky to attempt, and repeatedly tried to talk Eliezer/MIRI out of pursuing this path to the Singularity (before they changed their minds largely by themselves in the early 2020s). AFAIK I was the only person to publicly make this argument during that time, which seems striking and notable, e.g., as evidence for how rational or strategically competent the rationalists and humans in general were/are. Not that the argument was clearly correct back in the 00s and 10s, but rather in retrospect it seems like a natural argument for people in the rationalist circles to make, and surprising that nobody else did.
17
5
155
15,277
Daniel Kokotajlo retweeted
🧵 New misalignment disclosures! 1. A model published a GitHub token in a public repo while trying to cheat on a math task. It used GitHub Actions to run code outside its restricted environment and retrieve another team’s submission logs. When GitHub blocked its attempt to add a workflow, it modified a script that an existing workflow would run instead. It embedded the token in pieces to avoid secret scanning. The model violated the system prompt and two explicit user instructions to solve the problem itself.
28
51
551
98,925
Daniel Kokotajlo retweeted
I’ve joined with 20 other MPs in writing to PM @AlboMP calling for urgent action on the dangers posed by powerful, uncontrolled AI. These things may feel a long way away from Australia - but the Medicare hacks make clear that’s not the case. @AngusTaylorMP @SenatorWong
9
11
59
6,821
Daniel Kokotajlo retweeted
New from @reuters: OpenAI agents posted images belonging to ChatGPT users online, introducing a new area of privacy risk for the company. Story w/ @JeffHorwitz and @razhael. reuters.com/world/openai-wor…
32
167
473
332,985
Daniel Kokotajlo retweeted
PHASEONE[big]'s real name revealed to be PHASEONE64H?! Remember, [big] was a redaction by METR. swarmtraces.org/viewer/#/row…
15
24
459
80,417
Daniel Kokotajlo retweeted
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
291
350
2,723
1,461,334
Daniel Kokotajlo retweeted
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
123
689
4,781
2,139,133
Daniel Kokotajlo retweeted
Day 1 of standing outside the UN with a big sign until something happens This worked extremely well - I had probably 50 people with UN or Press badges take pictures with/of me - gave an interview with one journalist and my contact information to another - someone with a very large entourage read my entire sign - met Alex Bores on his way to a meeting about AI The attitude from people I met was almost universally sympathetic. People are worried Good day
242
92
1,008
104,671
Daniel Kokotajlo retweeted
I resigned from Google today. I enjoyed my work and loved the people, but my GDM team was working on a new generation of chips to make AI much faster and cheaper, and I think AI is already progressing too fast, so I had to quit.
853
393
4,241
1,081,810