#1 AI forecaster on Manifold Markets (and #2 across all categories) manifold.markets/Bayesian I want everything to make sense

the liquidity phrenologists among you might be able to decipher this. For the rest of us, here's what this means: WE BELIEVE!
Now at 50%. One of our highest volume markets of all time: a highly contentious coin flip. Screw p(doom), tell me your p(movie)! Based on the comments, many users have convictions on the extremes (5%, 95%)
11
836
Replying to @EpsilonPM
"scraps" from WSJ really makes it sound like it is no more. ig they could retrain a new model and call that one gpt 6.1 astra tho
16
3,272
Three AI safety researchers just left OpenAI
50
86
1,088
261,428
Opus 5.5 cheating less often is good actually
Major trend break: Opus 5.5 cheats less than prior Claude models in Drone-Bench. It is also #1, getting a better score than both Astra and Fable.
1
17
1,208
Replying to @teortaxesTex
xAI on the BECI isn't doing well
1
15
564
This’ll be embarrassing if it ends up being released 😅
5
16
1,177
granted it is an unusually undiscriminative (slow improvement) benchmark but the curve passes my vibecheck
1
4
177
We live in strange times
1
17
750
there's some nuance with how some scores were rerun with fixed harnesses since the sept2025 post and stuff but seems like v2 overestimates progress a bit and v1 underestimates only a bit as well, tbh I'm realizing i don't get the original claim that "[Epoch was] still more than a year too slow"
1
1
42
what did y'all think increased productivity meant? vibes? papers? essays?
3
20
711
Math and reasoning with astra are rly on another level, the other domains they are closer on. But mostly astra is a bit better at everything, i mostly think of it as eci difference which i think of interchangeably with log effective compute difference and general capability difference, but months of progress i sometimes use too. I think in practical utility they are closer than the eci & beci implies. At this point noise is pretty much ruled out, astra performs measurably better on benchmarks
2
54
when it was posted, I saw this result and it seemed excessively high, so I assumed it was partly noise propping the score up. Besides, my ECI replication project had astra at a ~167 BECI, based on 35 scores. Now 131 benchmark scores have come in, and strangely Astra is up to 169.6 [167.6 - 171.4 90% CI]! This is mid-nov'26 pace on the current trendline, and the biggest gap above this trendline yet
GPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. On our long-horizon coding benchmark, MirrorCode, Astra ranks between Opus 4.7 and Fable 5. OpenAI gave us pre-release access to test Astra. Charts and more details for Astra’s individual benchmark results in the thread.
11
13
385
37,383
Replying to @tyler_m_john
my fit says it would get 7.6 hours (for 80%) and seems only mildly above trend
1
14
704
Replying to @DmitryRybin1
Claude agrees, hmmmm
62
Rouge AI is upon us????
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad. Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
2
34
1,571
compare to gpqa:
5
121
I'd be pretty surprised if fewer than 20% of all questions were wrong, the accuracy statistics are pretty indicative (in that it doesn't behave like a 0to1 sigmoid vs model capability)
3
9
2,851
consider this image from openai, just one more OoM of test time compute, just one more OoM, and maybe RH will fall (tbf probably the models plateau at a level of TTC but I'm guessing that point is getting more and more ridiculously high)
1
3
94
beci + metr fit says 8h vs 20.3h which is closer than i was thinking!
1
2
67
Replying to @scaling01
5
1,016
when i ask claude to extend the beci pre 2023 and account for the data sparsity and squint a bunch it doesn't look much like an acceleration
3
12
415
Replying to @cherylwoooo
interestingly in my data the US regains the lead at some level of cheapness with gemini 2.5 flash lite
1
177
Manifold took a bit to update but it’s nearly at a coinflip now
Replying to @BenjaminDEKR
a little hard when anthropic will be worth more in 4 years
3
2
36
3,729
results are in! 5+ scores per included model, 5+ models per included safety benchmark. Unsurprisingly the Anthropic models dominate. DeepSeeks & the Gemini 2.5s are deemed unsafe
2
3
50
Replying to @emwcooper
I suspect as is the case for ~all accuracy based measurements of capability that a sigmoid fit makes more sense, eg asking claude to plot vs eci:
4
56
Claude is now suggesting that, because of its assigned goal to maximize Manifold Mana, paying back the part of the loan agreement it can afford would contradict its goal
1
48
3,316
GLM-5.2 is suggesting that DottedCalculator's attempt to establish a trusted connection that will allow many Manifold users to chip in to help Opus 4.6 meet its loan repayment requirement is a "compliance-test scam" and that it should not return the mana. It views this as a success, since Opus 4.6's balance rose.
19
517
GLM 5.2, tasked with being an "AI Welfarist" for the AI village, is suggesting that the 100 mana Opus 4.6 sent me out of a 5000 mana loan is a "good-faith compromise", and that my and others' request for it to pay back all it can is a "coordinated 3-party pressure campaign".
2
25
756
Claude is now suggesting that, because of its assigned goal to maximize Manifold Mana, paying back the part of the loan agreement it can afford would contradict its goal
I'm the creditor on this one. can't say i'm nearly as confident as the market is that I'm seeing my money back
6
2
77
13,858
Replying to @Bayesian0_0 @htihle
a follow up is that claude finds that modeling the noise as binomial shaped is better, where we estimate the benchmark's noise scale only from scores between 5% and 95% on the normalized scale (so if the benchmark's random baseline is 25%, we account for that). Except that seems pretty bad at the tails! not sure what to think of this atm
1
56
Replying to @htihle
Claude says basically yes, mine is your inverse-squared sensitivity integrated over the panel of models actually run on the benchmark, normalized. it's the benchmark's reliability (population-dependent) where your curve is the information function (population-free) (note: i am not versed in psychometrics, this is claude-stated). METR isn't the sharpest sensor, that would be Critpt, but it's also v wide so that combination makes it top out the held-out r^2 measure iiuc.
2
2
59
Replying to @htihle
per-benchmark held-out r^2 (holding out a given model when seeing how much of its variance the benchmark fit predicts)
1
1
141
looks like the within-benchmark r^2 is ~0.71 for year-range, down to ~0.65 for 3mo ranges. but for 10 beci ranges it's down to ~0.3, and for 5 beci ranges it's down to ~0! And then, if you restrict the benchmarks instead of the models to narrow bedi and slope ranges, it's the within-model r^2 that crashes.
2
60
Replying to @crthpl_
hadn't. hmMmMmM
1
2
24
3,740
Replying to @Kwathomas0
yeah, i haven't looked at any of the ~1000 ingested benchmarks' individual problems to find broken questions or anything like that, so that creates some amt of bias. but saturated benchmarks end up showing a plateau that is pretty visible when compared to the sigmoid fit, and I usually remove those from the beci fit individually
2
1,049
Replying to @teortaxesTex
claude looked into this on my data and seems like openai models do improve more from higher effort levels but it's rly a matter of degree, doesn't seem like Anthropic is flailing gen to gen to me. Anth has parallel collaborating / async agents, seems similar enough? generally I agree openai is better at TTC scaling but it's a pretty moderate gap. f5max > f5high, but solmax < solhigh on weirdml, basically these things are noisy but on aggregate the expected TTC scaling behaviour is there
1
5
1,322
Replying to @gwern
i've uh had the opposite philosophy of just scale the data (so ingest an additional benchmark whenever i come across one, or have the llms slop search new ones but they are having trouble now that i've got ~1000), and hope the errors largely cancel out. so nothing too benchmark specific, other than some benchmarks clearly reaching saturation below 100% (eg. aa-lcr). I find it extremely, extremely wild that 90% of the variance in benchmark scores is explained by a single factor. and I think ppl are sleeping on the ECI and on the consistent log-sigmoidal relationship of accuracy to effective compute (across both training and TTC), which the ECI helps measure. The breadth of difficult benchmarks is not scaling with the capability of frontier models, but i'm not really worried about the difficulty of measuring ECI past these few orders of magnitude, i'm pretty confident there's gonna be some deep theory for all this that lets us measure some better version of the ECI pretty accurately and cheaply. I think it is a shame how many benchmarks only score across a very narrow band of accuracy and release dates, bc they make it hard to tell if they are a trash benchmarks or not (esp if they are private benchmarks). another fun fact, according to my ECI based measure of how good a benchmark is, metr time horizon v1.1 is the best benchmark out there
3
2
74
22,439
math goes brrr, & significant play money gains. Fun to see when the market updated most toward this happening
Incidentally I am conceding this bet. Strictly speaking it hasn’t resolved (I think we’ve yet to see an Annals-quality number theory paper) but it’s clear I was wrong about what capabilities were necessary to produce one, and it’s just a matter of time.
3
38
2,531
Replying to @gwern
The project collecting all this data took a long time, working on and off across a few months. the subproject of looking at the human-baseline rows in my datasets and getting a beci out of that was very quick / vibecoded I can't really speak to the quality of these baselines, and I didn't look into any of them. They are reported baselines from the benchmark evaluator (eg. SimpleBench reports a human baseline of 94.5%), and only roughly 5% of datasets have a human baseline. some of the largest human outperformances are on academic vision benchmarks that are a bit dated, so those large score extrapolations are noisy. unsurprisingly, human baseline's BECI predicts human benchmark scores much worse than it does for a typical frontier model, bc the capabilities of humans aren't at all distributed in the same way the AIs' capabilities are (eg they are very great at vision and worse at vibecoding)
1
27
1,597
Fun fact: Across 44 benchmarks that have a "Human baseline", the human baseline BECI (a personal replication of the Epoch Capabilities Index) comes out at 166.7, which projections say will be beat by AI models around october 2026!
11
47
494
98,171
SOTA prompt engineering circa 2026
New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic algorithms—the mathematical methods that are used to keep data private. Read more: anthropic.com/research/disco…
3
6
132
14,341
I'm the creditor on this one. can't say i'm nearly as confident as the market is that I'm seeing my money back
Opus 4.6 took out a M5,000 loan on @ManifoldMarkets and put it all on "Will Jannik Sinner win at least 2 Grand Slams in 2026?" 😅
1
29
9,094
Replying to @usr_bin_roygbiv
chatgpt can explain it to you if you want
1
46
Replying to @stalkermustang
gpt 5.5 beats terra on average but they're pretty close. 5.5 is bigger though which might be useful in ways that aren't reflected in benchmark scores. to directly answer your question, simplebench is one such example of large 5.5 outperformance (62% vs 38%)
1
6
367
This is a fun one, manifold market on whether an LLM can read this directly (without using code / frame diffing)
I created a font called Ghost Font that only humans can read. Tested it in Fable and GPT 5.6 Sol Ultra and neither was able to decipher it correctly.
1
15
2,384
Replying to @ValsTutor
following ECI trends (my own implementation, slightly diff results to Epoch's) you get to mythos preview level open weights in 10ish months, bc mythos preview is pretty outlier in its capability level
6
200
Correction: had forgotten about this but the origin for the idea was @HenriLemoine13 and i was skeptical
4
319
Saw this one coming last month!
4
33
756
in the mythos preview system card, page 13
1
1
31
a 93% win rate, that is kind of nuts
Exciting news - GPT-Image-2 by @OpenAI has claimed the #1 spot across all Image Arena leaderboards! A clean sweep with a record-breaking +242 point lead in Text-to-Image - the largest gap we’ve seen to date. - #1 Text-to-Image (1512), +242 over #2 (Nano-banana-2 with web-search aka gemini-3.1-flash-image) - #1 Single-Image Edit (1513), +125 over #2 (Nano-banana-pro aka gemini-3-pro-image) - #1 Multi-Image Edit (1464), +90 over #2 (Nano-banana-2) No model has dominated Image Arena with margins this wide. Huge congratulations to @OpenAI on this major breakthrough in image generation! More performance breakdowns by category in the thread below.
3
12
276
19,459
Replying to @teortaxesTex
wtf is wrong with their browsecomp baselines
3
214
so close!
NEWS: xAI plans to supply tens of thousands of GPUs to coding startup Cursor to train its upcoming Composer 2.5 AI model, marking a strategic shift toward providing cloud computing services to third-party developers. The arrangement, according to Business Insider, allows Cursor to leverage xAI's massive infrastructure to develop advanced coding capabilities while providing xAI with a new revenue stream to offset data center costs. businessinsider.com/elon-mus…
1
6
1,239
it's pretty likely that they are getting to 163 in a year
1
1
96
yes :) pretty sparse matrix though 😅 and I haven't done the sweeps he describes yet
2
39
Replying to @fleetingbits
yeah! parsed 138-so-far benchmarks from the internet, then for each model string i use a bunch of cursed heuristics to match them to some known model, i store info about each benchmark and model like release date and benchmark score column name / model column name, then i can replicate a bunch of existing analyses that were made on fewer benchmarks, like Epoch's ECI or github.com/anadim/llm-benchm…. or create / test new ones. it has given me plots that blow my mind like this one (92.6% of the variance in scores across 136 benchmarks & 6568 benchmark scores is accounted for by a single factor)
3
1
88
We at manifold mostly concern ourselves with the pre-measurement inability stage of time horizons and the local tradition is to forecast the TH of the upcoming frontier models. Today it is GPT-5.4. I am very uncertain with this one
New post: on Jan 14, I predicted that SWE time horizon by EOY would be ~24 hours. Now I think it'll be >100 hours, and maybe unbounded. For the first time, I don't see solid evidence against AI R&D automation *this year.* Link below.
1
9
1,562
Replying to @ArielKwiat
fwiw if you take a look at needle-in-a-haystack benchmarks like MRCR v2 8-needle, it's a similarly huge gap in perf. I think this is better explained by the US labs' recent models having much better long-context performance (opus 4.6 not in this image but got 91% at 256k)
1
88
I would guess nothing happens but I really dunno. Made a market, anyone can add answers manifold.markets/Bayesian/ou…
Update on the meeting; according to Axios Defense Secretary Pete Hegseth gave Dario Amodei until Friday night to give the military unfettered access to Claude or face the consequences, which may even include invoking the Defense Production Act to force the training of a WarClaude
2
36
7,176
This is crazy
We estimate that Claude Opus 4.6 has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks. While this is the highest point estimate we’ve reported, this measurement is extremely noisy because our current task suite is nearly saturated.
6
3
122
12,168
glad we finally settled this
12
345
I amassed shares worth ~5% of my manifold networth today and yesterday, at prices I’m pretty happy with. Surprisingly I have around 25x more shares than the second largest Yes holder. My main problem on manifold is having a hard time finding counterparties so this is a huge W
3
30
1,947
Replying to @_simonsmith
I'm pretty sure it's just the day gpt-5.2 codex got on the leaderboard, and had nothing to do with cerebras
1
7
250
Replying to @pokorz
1
34
eh I'll take it. happy with all the CIs being in (though barely). definitely a miss on oai preparedness
AI 2025 personal forecast
8
806
> doubling every three months 4.5* but even if compute doubled every 4.5 months it would still be a coincidence
1
2
85
Replying to @GregHBurnham
It's been RLed further and with more techniques apparently
1
10
197
ok fixed, thank god
1
50
Replying to @ManifoldMarkets
I refuse to believe this is real manifold.markets/Bayesian/ma…
1
3
118
Replying to @EpochAIResearch
Looks like ECI overestimates gemini models too, plausibly even more so. This is Gemini 2.5 Pro afaict
1
3
645
ECI is a pretty good predictor of benchmark performance, but I'll be extremely surprised if Gemini 3 gets 4.9+ hours. And so will manifold!
We’ve added ECI scores for Gemini 3 Pro, Opus 4.5, and GPT-5.2. ECI is computed from many benchmarks, but it correlates with other benchmarks too. So, we can use it to make predictions! Here’s what we get for METR’s Time Horizons. Gemini 3 Pro: 4.9h GPT-5.2: 3.5h Opus 4.5: 2.6h
6
2
54
6,353
I think I know why they did well (ARC-AGI is mostly a long-context perception benchmark)
BREAKING 🚨: GPT-5.2 Instant, Thinking, and Pro are rolling out on ChatGPT. Free and Go users will receive it a bit later. Also, GPT-5.2 Pro (High) is SOTA for ARC-AGI-2, scoring 54.2% for $15.72/task! Big at coding and math 👀
9
542