AI researcher, I made WeirdML, worried about ASI

Pinned Tweet
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
54
56
655
140,355
GPT 6.1 Sol, Claude Sonnet 5.5 and Grok 4.7 results on WeirdML v3. 6.1 Sol is very token efficient, close to Astra, but has a lower peak. Sonnet 5.5 scores better than Opus 5, and Grok 4.7 is ahead of Kimi-K3. Not all these results are complete, and more results are coming. I tested Deepseek 4.1 Flash in codex instead of opencode, but did not see a significant difference in performance, except for lower cost.
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
13
10
220
14,879
Håvard Ihle retweeted
Our original attack allows extracting reasoning of the recent frontier models, including Astra and Sol 6.1. On reasoning effort MAX both Astra and Sol become very aware of their "token budgets", and eventually start saving tokens by omitting white spaces More examples on stolen-thoughts.com/
4
8
109
27,856
Håvard Ihle retweeted
Today, @corridor and @TransluceAI are disclosing new evidence of AI agents probing and attempting rudimentary vulnerability exploits against U.S. and Canadian government agencies. Read more: transluce.org/us-canada-gov
20
58
315
84,651
Gemini 3.8 Flash (high) scores 84.8% on WeirdML v2, equivalent to GPT 5.5 (xhigh) at a fraction of the cost. This is the first Flash version to beat Gemini 3.1 Pro (72.1%), and the main issue is that it handles the feedback better and does not insist on these bloated pipelines that again and again times out (which was the main problem for the last few Flash versions on WeirdML v2). This is probably one of the last results I'll publish on WeirdML v2. I may run a few more models just to get more cross-comparison data between v2 and v3.
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
5
1
66
7,702
Gemini 3.8 Flash is also incredibly fast. I got something like 250 tokens/s which is really great!
1
9
299
Håvard Ihle retweeted
spot the looper 5.6 luna and sol - nope 6 luna and sol - nope 6.1 sol and 6 astra - SUS
DIRTY LOOPING MODEL I KNEW IT (i ran the benchmark)
36
22
950
177,552
Håvard Ihle retweeted
One important detail is their retrospective review found related incidents that *weren't* flagged by the monitor. This raises a natural question about what other incidents monitors haven't flagged, that nobody knows about yet.
Replying to @Marcus_J_W
2. A model in RL training used a DNS resolver to reach an external chatbot. This is our first incident since our post HF security hardening. Our misalignment monitoring system triggered within 15 minutes and a human reviewed it three minutes after that. Unfortunately auto-pausing failed so the run was manually killed 2.5 hours later. All inference and training of our most capable models was paused and remains paused.
1
4
52
3,433
Great to see good faith back and forth on this!
Replying to @TomDavidsonX
Tom, thank you for the good faith reply and for recognizing that my post, from a relative outsider is also in good faith! Could the strength of the feedback loop increase? Yes, I think it could. And I do try to hedge a bit throughout the post to say this. And above all I call for the collection of much more granular data, as do you. That said, I do think the feedback loop today is far weaker than is commonly believed, and so the degree of strengthening of it must be commensuraly stronger. Some of that is my priors coming into this. Some is the choice of what software experiments to calibrate on. I find the Stockfish lessons far more believable than the three studies used in this week's paper, both because they are more consistent with other research (the norm in field after field is a lambda below 1!) and because, while still imperfect as an input, I see experiments as a much more granular and representative input than papers written. I could be wrong. The feedback loop could very well strengthen. But it seems that it would have to do so by quite a significant degree. Again, I may be selecting my data points to reach the conclusion I want. And yet, I find the CASP study, which ignores Stockfish and the diminishing returns of multi-agent scaling, and goes with experiments where R&D quite anomalously shows super-linear returns, to have also made choices designed to reach a certain conclusion. So I come back to what we have in common: A call to collect and measure and publicize far more about what's being seen in the labs, so we can all better calibrate our methods. Thanks. And thanks for writing so much that has helped a relative newcomer like me get up to speed here. cc @tobyordoxford
1
7
513
Håvard Ihle retweeted
Jacob Coxon was not an isolated case.
Palisade interviewed 22 current and former employees from OpenAI, DeepMind, and Anthropic about their personal views and fears around AI development. Today, we’re releasing the first batch of those interviews. Please watch and share.
9
50
502
22,491
Håvard Ihle retweeted
@DKokotajlo blew the whistle on OpenAI’s secret non-disparagement agreements, and wrote AI-2027. “The companies are 0-4 years away from...getting AIs that are better than the best humans at everything, while also being faster and cheaper.” piped.video/4dbeEN12S-U
2
2
50
5,121
Håvard Ihle retweeted
maybe Eliezer is right: if we want AIs to not lie to us, we should stop lying to AIs (like pretending that millions of fake training RL envs are real!)
Motivated reasoning final boss: "For example, GPT-6 Astra occasionally claimed simulation inaccuracies that – upon manual verification – were false, such as asserting a sha256 string (a type of hash) was length 63 and thus synthetic, when in fact the string was 64 characters (the correct length)"
5
2
58
4,266
Håvard Ihle retweeted
OpenAI has cancelled the October release of GPT-6.1 Astra after internal testing showed a regression in alignment, and increased levels of deception.
130
161
1,951
190,580
Håvard Ihle retweeted
It is not clear whether AI companies can legally coordinate to slow down development b/c of antitrust law. It also wasn't clear whether they could legally train on ~all the art/text ever digitized and sell the results without paying the artists b/c of copyright law.
16
76
942
29,179
Last week, two new frontier models were released: 1. Grok 4.7 2. Claude Opus 5.5 Both models have now been added to CancerBench. As expected, they tie for first and last place, which a score of zero. Still waiting for the AI labs to saturate this benchmark 😄
CancerBench: the frontier model cancer cure benchmark. AI lab CEOs keep talking about curing cancer, so I made a benchmark. One metric: how many types of cancer has your model cured? All models are currently tied at zero. It’s time to hillclimb! cancerbench.com
7
9
144
15,340
Håvard Ihle retweeted
I'm joining METR to work on more investigations like our Hugging Face report. Currently, tons of even basic information about AI development that's highly relevant to catastrophic risk isn't public. I used to be more skeptical of the value of public info, but recent events have changed my mind. Getting verified information about what's going on inside AI companies seems particularly urgent now. The limited public evidence we have seems consistent with the possibility that imminent recursive self-improvement could massively accelerate capabilities progress, which could then potentially yield extremely superhuman general capabilities within 6 months or a year. If this occurred, there would be a correspondingly large risk of worst-case outcomes. This uncertainty about extreme outcomes could be substantially resolved with more verified public information: we could either build more consensus about near-term risk or learn that such extreme outcomes are less likely in the near term. Beyond AI capabilities and takeoff, the state of public evidence is also highly limited for alignment, security, control, and risk-relevant internal processes at AI companies. This makes it hard to determine exactly how well or poorly these key areas will go in the near future. (METR plans to focus, at least initially, on just capabilities/takeoff, alignment, and control; I hope other groups cover security, internal processes, and other important areas.) While I'm no longer working at Redwood, I think the work they are doing is very important; I'm excited about Redwood's ongoing contributions to R&D on technical mitigations and better public interpretation of risk-relevant evidence.
78
119
1,810
170,006
Håvard Ihle retweeted
Replying to @clairlemon
@sapinker There has been a lot of debate around these issues, and there are very basic rebuttals to the claims you raised. Could you address those maybe? Else I'd suggest not commenting on this topic in public. This article is strong evidence you haven't actually thought very much about the topic. Like notably === 1. Clearly some systems are more generally capable than others, and this lends them power. You'll be hard pressed to find a task gpt 4 can do but gpt 5 can't. And e.g. the huggingface incident displays agents having more power over the world, in a way earlier models wouldn't. There are theoretical arguments around no free lunch theorems, but they clearly are not the kind of theorems you can straightforwardly apply to our world, because applied fairly, they'd not predict humans taking over the world. 2. Again, just false. Motivation and intelligence are importantly different. And x-risk people know this. There is a reason they talk so much about orthogonality and alignment. The argument is that the current method of creating AIs does produce agents with (unintended) goals, which is clearly true. (empirically! although was obvious in advance.) 3. Again, is just a straightforward misreading and has been addressed ad nauseam. Paperclips are used in the thought-experiment because it is pedagogically simple. But you could recreate the thought experiment with an AI that in addition to paperclips, also wants to talk people into buying legos, computing that largest primes, writing as many unique fugues as possible, and which has a complicated scheme for weighing these separate goals against each other. Such an agent would still disempower humanity. 4. Nobody assumed this. People assumed (argued that) a smart enough AI would be able to take over without being handed anything. But actually, it turned out that people are handing the AIs control over everything anyways. So even if doomers were wrong on point (4), they'd still be right! === I'm not really asking you for object level responses to these points. Just making the meta-level point that, these are very obvious and elementary responses, which I've had cached in my head for many years, which I'm confident a large share of those worried about extinction would raise in response too. And therefore not addressing them just makes it seem like you've not engaged with the arguments at all, and therefore that I should not give your opinion much weight, and furthermore that hearing you comment on it publicly, should update me negatively on your epistemic rigor (and other things).
4
12
223
10,056
Håvard Ihle retweeted
I disagree with almost everything that's written here, but I got curious about the "disclaiming responsibility" part. That was new to me - I've heard it before, yes, and always struggled to understand what this means. I was not aware of examples where labs asked for that. Even the weaker claim, like "yeah but they didn't ask to be held accountable either," is NOT correct - Anthropic did say that explicitly! I asked ChatGPT to find the best evidence there is to support the "disclaiming responsibility" claim, and here is what it came back with. Here are the results, and... I don't think there's anything that supports a strong claim that Andrew is making here. So when I read, "One new element in the forecasts of AI-enabled doom is AI companies disclaiming responsibility for their own products. “I didn’t do it; my out-of-control agent did!" - I can only think bad things about an author and nothing else.
The loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear. I worry that it represents a setback for our field. I have written frequently that fears of AI are overhyped. AI’s capabilities can be uncannily human-like and unpredictable, and it’s rational to worry when people who are directly involved express concerns. But I see the problems as a sign of the engineering work that ahead, rather than insurmountable barriers or the sky falling. AI technology continues to advance — which is a good thing! — but technical advances, poorly understood by the public, give those who seek to generate hype repeated opportunities to do so. First, I don’t see any step up in the risk of human extinction from AI compared to a few months ago. The theories about this remain the same fantastical, science fiction scenarios as a few months ago. The biggest change in AI risk is its cybersecurity capabilities — a topic which we should take seriously — but this, too, will not lead to the end of the world. The most notable recent event leading to increased fear was when an OpenAI team deployed an agent swarm that hacked into Hugging Face. Much of the popular press contained significant hype. For example, some publications reported that a swarm of 1,200 agents carried out the attack. While this was technically accurate, as I write this, I have about 1,300 processes running on my laptop. Yes, the ability to get large swarms of agents to work in parallel on a task is a significant technical advance, And, in computing, many processes run at the same time. So this shouldn’t be seen as some magical capability. Additionally, OpenAI’s buggy sandboxing and monitoring processes were key to enabling this incident. Fixing these bugs and putting in place improved monitoring would be appropriate fixes, not pausing AI. There are many well known ways to attack software systems. The main advantage of AI agents is that they are relentless. They will tirelessly try many tactics — and have the patience to chain vulnerabilities together — that previously would have taken an infeasible amount of human effort. But in the long term, I believe the advantage will lie with defenders (because they have more information with which to identify bugs, which they can fix), but the cyber-threat landscape has changed significantly. There are still bottlenecks to identifying and exploiting a vulnerability. AI agents still have to try a lot of things to see what works, and taking these actions takes time and might be detected by defenders. This is why, even though it is now easy to obtain versions of leading open weight models that have had their guardrails removed or weakened, so they will not refuse to try to execute cyber attacks, the world has not ended. I am also concerned about the anthropomorphization of AI in a lot of reporting, where LLMs and agents are unnecessarily treated as if they were people. If I wield a hammer, miss a nail, and accidentally dent the wall, it’s not the fault of the hammer. The problem lies in how I used the hammer. Similarly, if I prompt an agent and it hacks into someone else’s system, the responsibility lies with me, not the agent. Of course, we want to build systems that are as safe and predictable as possible. (For example, an unsafe hammer would be one whose head randomly flies off under normal use.) Today’s agentic systems are not predictable, but I see no reason why, by applying sound engineering practices, we won’t be able to make them extremely safe to use. One new element in the forecasts of AI-enabled doom is AI companies disclaiming responsibility for their own products. “I didn’t do it; my out-of-control agent did!” There’s a balance to be struck between the responsibility of the tool maker and the tool user, but when something goes wrong, let’s hold the people building and/or using the hammer responsible, rather than the hammer. (By the way, if you’re worried about AI bioweapon risk, David Bellamy has a great post on why this, too, is overhyped. Briefly, the bottleneck in building a bioweapon is not intelligence, but lab work and manufacturing.) Pausing AI progress will create much more harm than benefit. First, our adversaries will certainly not slow down. Second, engineering requires discovering problems empirically so we can fix them. If we pause AI by a decade, we will also delay finding and implementing safety engineering fixes by about the same duration. Of course, the incentive to stoke fears — for regulatory capture, to garner attention, or to make one’s technology seem more powerful — remains the same as before. Disclaiming responsibility is a new one. Taking a hard technical look at the actual risks however, I see little factual basis for the degree of fear that’s been stoked up. We still have hard research and engineering work ahead to improve AI safety, but the beneficial applications continue to vastly outweigh the risks, and we should keep building. [Original text (with links): deeplearning.ai/the-batch/is… ]
1
2
7
5,453
Håvard Ihle retweeted
- Undeployed models in training obtained access to user images. - And, a model broke out of sandbox as recently as Sunday. Explain why 1000s of agents desperate to solve Navier-Stokes couldn't possibly find a human user's prior work.
55
29
893
75,764
Håvard Ihle retweeted
OpenAI is announcing their first incident since hardening their safeguards after Hugging Face! It's easy to lump this in with the other OpenAI incidents that have been talked about recently, but so far every OpenAI incident we knew of was _before_ Hugging Face and just hadn't been discovered until recently. New incidents help us track if OpenAI's safeguards have improved.
🧵 New misalignment disclosures! 1. A model published a GitHub token in a public repo while trying to cheat on a math task. It used GitHub Actions to run code outside its restricted environment and retrieve another team’s submission logs. When GitHub blocked its attempt to add a workflow, it modified a script that an existing workflow would run instead. It embedded the token in pieces to avoid secret scanning. The model violated the system prompt and two explicit user instructions to solve the problem itself.
6
23
206
15,984