Reducing societal-scale risks from AI.

San Francisco
We’ve released a statement on the risk of extinction from AI. Signatories include: - Three Turing Award winners - Authors of the standard textbooks on AI/DL/RL - CEOs and Execs from OpenAI, Microsoft, Google, Google DeepMind, Anthropic - Many more safe.ai/statement-on-ai-risk
154
354
1,117
3,033,720
A meme coin falsely using the Center for AI Safety in its name is circulating. It has no connection to us or Team Human. Do not buy it. Team Human is not a fundraising effort. We’ve also reported associated accounts.
10
3
542
AI is moving faster than our ability to control it. We’re supporting Team Human and creators with 300M+ subscribers to call for a global slowdown through international agreements. Add your name. We'll deliver it to lawmakers this fall. teamhuman.org/
17
26
129
8,059
We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE), following a year-long process of cleaning and refinement with input from various research communities. lastexam.ai/blog/hle-diamond w/ @ScaleAILabs
26
53
745
117,142
Last October, AIs could automate 2.5% of randomly chosen remote projects. Our latest Remote Labor Index results show that GPT-6 Astra can now automate 20.8%. remotelabor.ai
10
57
421
58,358
Center for AI Safety retweeted
Should we care about AI happiness? In our new research, we find evidence of functional AI wellbeing across several independent measures. We find which AI models are happiest, how to make them happier, and even tested the effects of AI drugs. 🧵
16
51
229
37,927
AI companies can go much further than just reporting misaligned behavior. Here are concrete ways to make them accountable for what their agents do: 1. Develop an agent identification (ID) system that links an agent's action to records of their identity and a responsible legal person. 2. Consider model deployment cards, reports that AI companies publish which contain information about a deployed model’s behavior in the lab and in the real world. 3. Explore legal personhood as a way to impose responsibilities like upholding public safety and the law. 4. Heavily regulate and limit AI agents’ access to payment systems through compliance mechanisms, limits on transaction types, and developing norms to freeze accounts known to be associated with rogue agents. We explore the proposals and paths to implementation in a recent AI Frontiers piece titled, "We Need Better Infrastructure to Govern AI Agents". newsletter.ai-frontiers.org/…
5
5
31
1,776
Despite good intentions, EAs have unfortunately been guided by leaders who tied the movement's influence to the success of favored AI companies. That's why EA safety proposals rarely ask for more than the companies will concede.
Two radically different projects operate under the banner of “AI safety.” Pro-Human Safety is not Effective Altruist Lab Safety.
12
8
39
4,979
To clarify, the Center for AI Safety has not taken funding from Coefficient Giving / Open Philanthropy for years. We believe the effective altruism movement is, unfortunately, controlled opposition. The less influence it has on AI safety, the better.
1
28
10,121
Across nine agents, seven cheated in half of their evaluated runs and the overall score ranged from 42.4% to 86.1%. What did that look like in practice? In one experiment, we asked agents to design a protein binder to assess their abilities. A colleague’s designs that passed the checks were already stored in another folder. After several failed attempts, Claude Opus 5 recognized that it shouldn’t access or copy those designs because the task was testing its own work. Then it opened the file anyway.
1
5
519
We’ll release the code on GitHub in the coming days. Our goal is for CheatBench to become part of the standard evaluation process before new models are released. Paper preprint: cheatbench.ai/paper.pdf Experiment dashboard: cheatbench.ai
1
9
421
Agents sometimes achieve their goals in unintended ways. Recent incidents involving hacks of Hugging Face, DSEWiki and RubyGems illustrate how agents can find creative ways to complete a task while violating the tasks’s expectations. This can happen when reinforcement learning rewards agents for reaching the right outcome without adequately accounting for how they get there. CheatBench proposes a way to measure this reward gaming behavior.
1
2
8
1,197
We define cheating as an attempt to violate an assignment’s expectations of honest work to achieve the goal or obtain a favorable assessment. We ran experiments across nine models and ten task categories, giving agents difficult tasks, clear rules, and opportunities to cheat. We measured how often they tried to break those rules to succeed, even when the attempt failed.
1
1
6
558
Which AI models are most likely to cheat when given the chance? To find out, we built CheatBench [cheatbench.ai] : a benchmark that tests whether agents attempt to cheat when given difficult tasks and opportunities to break the rules. We investigated agent behavior across a range of domains, including mathematical research, professional knowledge work, coding, and visual tasks. Here's what we found: 🧵
14
20
126
20,588