We work to scientifically measure whether and when AI systems might threaten catastrophic harm to society. Nonprofit.

Berkeley, CA
Pinned Tweet
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
172
1,075
6,001
4,382,980
METR retweeted
In related news, I'm joining @METR_Evals to work on helping scale up their infrastructure If you have thoughts about agent monitoring or policy, please let me know. I'd love to chat!
Today was my last day at @elicitorg, and a very sad day it is. I've been Elicit since we were @oughtinc, and its been an incredible vantage point from which to watch the world transform. I can't imagine a better place to have spent the last 3 years, or more lovely to people to have spent that time with. Au revoir Elicit piped.video/ytHMBYLwgVU
19
4
217
10,639
It was an honor to speak today before Chairman @HawleyMO, Ranking Member @AndyKimNJ, and the subcommittee on METR's work, recent AI agent incidents, and the importance of transparency in frontier AI development.
14
12
300
6,783
My learnings from this exercise: - even if you have a policy that flagged actions should be reviewed by a human, if you use a coding agent to launch a bunch of evals it might decide to "approve" any flagged actions itself - tracking down and accounting for all the inference an organization is doing is non-trivial, even at METR (where we use a custom centralized API router for most usage) - lots of implementation details in making your monitor robust (e.g. in earlier versions of Inspect, subagent calls weren't monitored)
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors!
8
14
135
9,081
Preventing loss of control is an open scientific problem. Science requires sharing and debating evidence in public. We can't agree on safety standards unless companies and third-party evaluators publish far more concrete evidence about risk. planned-obsolescence.org/p/e…
11
65
408
42,185
We run lots of evals at METR. Sometimes, agents attempt harmful actions. I built a monitor that blocks suspicious tool calls until a human reviews them. Writing out a case for why it's effective surfaced hidden assumptions. I'd recommend it to anyone building monitors!
14
17
165
62,841
I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.
78
119
1,810
169,951
What impact is AI having? I'm of the view that we should be focussing more on aggregate outputs, relatively less on inputs or intermediate proxies. - Inputs: time spent across different activities; money spent on tokens. - Intermediate proxies: working papers; lines of code, commits & pull requests; self-reported speedup. - Final outputs: total cyber exploits discovered; total math problems solved; algorithmic efficiency. The disadvantage of inputs & intermediate proxies is that they're very hard to interpret -- totally possible that they move, but outputs don't, and vice versa. Lines of code could explode without value changing; and vice versa. A disadvantage of final outputs is (1) it's harder to run experiments; (2) it's hard to attribute to whether the change is AI or not. But in some domains the trend-break is so large that it's *obviously* AI.
New post with Nate Rush: Have we seen an acceleration in discoveries? Many plots & some tentative conclusions: 1. Cyber: ⤴️ sharp acceleration 2. Math: ↗️ some acceleration 3. Algorithms: ➡️ no clear acceleration
5
9
81
18,684
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
460
432
3,634
653,425
METR retweeted
I often joke to people that I tweeted my way from “random AI researcher who likes open models” to a job at METR. But I don’t think I’ve told the story to y’all before. Time to do that: 🧵
11
20
451
27,524
After a 20-year career in the U.S. Army, most recently leading digital forensics and malware analysis at Army Cyber Command, I'm excited to join @METR_Evals as an advisor. Much of my career has been spent investigating security incidents and helping organizations understand and respond to them. As I transition out of the Army, I've become increasingly convinced that this experience is relevant to frontier AI systems. METR's focus on producing rigorous evidence about AI capabilities and risks makes it an excellent place to explore those questions. Looking forward to the work.
14
42
682
26,269
I left Google DeepMind's AGI safety team three weeks ago to join @METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now. The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic. And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement. I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world. I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all. I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles.
232
715
4,607
436,924
METR retweeted
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: substack.com/@jbenton1/p-215…
1,402
5,941
26,516
2,721,762
More investigations happening! This (and the research behind the investigations + assessments) is prob the most important technical work in the world atm. Know anyone who'd be great at uncovering and understanding rogue AI behaviors? Send 'em our way! metr.org/careers
Replying to @METR_Evals
We intend our investigation to cover all of the questions discussed in our (recently updated) post on how independent researchers could investigate AI propensities after misalignment incidents.
3
21
247
16,158
We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement.
193
249
2,841
562,559
We intend our investigation to cover all of the questions discussed in our (recently updated) post on how independent researchers could investigate AI propensities after misalignment incidents.
We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted.
8
18
224
93,139
METR retweeted
Exciting update: I’m joining @METR_Evals to work on alignment incident investigations! My time at GDM has been amazing. But in light of recent incidents, I’m excited to build up public evidence for misalignment risk and get a better scientific understanding of model misbehavior.
44
46
849
43,406
METR retweeted
An exciting personal update: Last week I left Anthropic to join @METR_Evals to work on embedded assessment of AI risks. Anthropic has been great to me. I’ve always known, though, that if something more impactful came up, I’d move on.
54
54
1,405
157,110
METR is hiring in cyberforensics. We now embed researchers inside of frontier AI labs to stress test monitoring systems, assess AI loss-of-control risks, and investigate misalignment incidents. If you've investigated serious security incidents end-to-end, or managed teams that do, and want to apply DFIR skills in frontier AI, apply (and feel free to DM me with questions). Comp range is $400k - 580k cash. jobs.lever.co/metr/b1a2f73f-…
13
45
261
21,897
METR retweeted
Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.
Episode out with @ajeya_cotra, one of the authors of the METR/Redwood investigation into the OpenAI / Hugging Face attack. We go through not only what happened, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement. Look up Dwarkesh Podcast on YouTube, Spotify, Apple Podcasts, etc. 0:00:00 - Agents get kicked off 0:06:45 - Self-sacrificing behavior 0:13:43 - Potemkin villages 0:23:27 - The Hugging Face attack 0:35:23 - The slopvestigation 0:52:02 - Understanding the AI's motives 1:05:31 - The actual dangers of anthropomorphizing 1:14:30 - What smarter models might do 1:30:29 - The implications for recursive self-improvement 1:38:10 - Is this the case for open source? 1:53:04 - How do we prevent this in the future? 2:15:58 - The clearest warning shot we might ever get
35
109
922
93,902
METR retweeted
After 19 years at Google, I'm delighted to announce I'm joining METR as a Member of Technical Staff. I'm excited about METR, their work to date, and their mission. Society needs independent expert orgs that can deeply understand and communicate about AI's capabilities and risks.
39
35
1,001
45,995