#1 AI forecaster on Manifold Markets (and #2 across all categories) manifold.markets/Bayesian I want everything to make sense

I think AI is probably conscious despite occasionally touching grass
AI isn’t conscious. If you think it is conscious or will be conscious, I personally think you need to remove any access you have to model weights asap. Now step away from the laptop or keyboard, put down your devices and go outside and go touch some grass.
1
11
434
Bayesian retweeted
this is astras answer on what it did 😭
2
2
35
1,524
the liquidity phrenologists among you might be able to decipher this. For the rest of us, here's what this means: WE BELIEVE!
Now at 50%. One of our highest volume markets of all time: a highly contentious coin flip. Screw p(doom), tell me your p(movie)! Based on the comments, many users have convictions on the extremes (5%, 95%)
11
827
Three AI safety researchers just left OpenAI
50
86
1,087
260,952
💯 "I'm not convinced that racing to RSI is actually a prisoners dilemma - I think it might actually just be insane and hubristic in the way it naively appears to normal people such that additional unilateral restraint is in fact rational and moral."
1
1
49
7,802
Update:
Wow, per the WSJ these safety researchers were fired for 'allegedly sharing confidential company information with a third-party AI-safety organization'
2
4
65
6,807
Dario Amodei's wife just left Anthropic??
👀 Cami Clark has left Anthropic.
28
17
827
181,728
I’m quite unhappy with much of what OpenAI does. I am very happy that I’m allowed to say “I’m quite unhappy with much of what OpenAI does.”
1
65
Opus 5.5 cheating less often is good actually
Major trend break: Opus 5.5 cheats less than prior Claude models in Drone-Bench. It is also #1, getting a better score than both Astra and Fable.
1
17
1,207
this is twitter -> X for AI people
President Trump 'We're going to be signing a document today at about five o'clock, renaming Artificial Intelligence, because it's not artificial, we all agree on that, and we're going to be renaming it Super Intelligence. Officially renaming it.'
2
2
30
2,092
Bayesian retweeted
opus-5.5 made a ytp with darios CBS interview, actually quite funny in some places audio on
37
78
1,266
69,942
Bayesian retweeted
Based on my estimates, the 10,000-agent swarm that solved the Navier-Stokes problem was about 8 months of frontier progress ahead of a single GPT-6-Astra instance (with a range of 5.5 to 12.4 months) It used 48,000 times more output tokens than the largest single model evaluation and roughly 68 times more than an entire long-horizon coding benchmark
Just dropped my new article. Hope you enjoy it :) scaling01.substack.com/p/acc…
35
97
1,518
179,622
Bayesian retweeted
Introducing WeirdML v3, a fully agentic benchmark featuring 11 complex hand-made tasks. Models must explore and understand unfamiliar data, develop ML and data analysis pipelines and produce results despite limited data, unspecified goals and/or very limited feedback. 1/8
WeirdML v2 is now out! The update includes a bunch of new tasks (now 19 tasks total, up from 6), and results from all the latest models. We now also track api costs and other metadata which give more insight into the different models. The new results are shown in these two figures. The first one shows an overview of the overall results as well as the results on individual tasks, in addition to various metadata. The second figure shows cost vs performance and shows a clear scaling with better results for higher costs. We also have a very varied pareto frontier with 11 models from 6 different companies having the best accuracy for a given cost for at least some of the cost range. Grok 3, Claude Opus 4 and GPT 4.5 are the ones that underperform for their costs, while Gemini pro and o3 pro have the best results at the highest costs. Qwen3 30B3A, grok 3 mini and deepseek R1 also each represent a good chunk of the pareto frontier.
54
56
655
140,262
6 months ago: Dario Amodei meets Susie Wiles and Scott Bessent at the White House about Mythos. Asked about Amodei's visit, Trump says "Who?" cnbc.com/2026/04/17/anthropi…
1
11
224
8,055