Research at OpenAI. Be kind to others, and yourself.

San Francisco, CA
Ted Sanders retweeted
wowi GPT-6.1 Sol scores 100% at FrontierMeth Tier 4 v2; shoutout to @AcerFur xoxo
19
15
432
20,798
Ted Sanders retweeted
I used GPT-6 Astra to break an unsolved cipher to one of Napoleon's generals that had gone unread for 217 years. What makes this impressive isn't actually the codebreaking, but that Astra completed the entire multi-modal workflow in ~6 hours from a single image and goal. 1/6🧵
313
1,612
11,707
5,255,143
One cool thing about ARC-AGI-3 is that it measures cost until success or timeout, instead of cost of first attempt. This is arguably more aligned with real life, and shows how costly models at high efforts can counterintuitively be cheaper, by needing less iteration.
GPT-6.1 Sol from @OpenAI on ARC-AGI (Verified): - ARC-AGI-3: 52.7%, $7.6K (standard harness), 96.4%, $4.4K (provider adapter harness) - ARC-AGI-2: 94.2%, $0.25/task - ARC-AGI-1: 98.5%, $0.06/task Its 96.4% on v3 was comparable to GPT-6 Astra's 99.9% but at a 77% lower cost.
1
6
76
6,728
Ted Sanders retweeted
GPT-6.1 Sol is exceptionally good at complex knowledge work tasks On ClickUp Bench, our internal eval for white collar knowledge work, it expands the pareto frontier with superior performance to models 2x in cost (!) Congrats to @OpenAI on a banger launch. Live in ClickUp now!
2
17
151
17,193
Ted Sanders retweeted
GPT-6.1 Sol is the new #1 on MathArena! Cheaper and better than Astra.
21
52
520
27,589
Ted Sanders retweeted
gpt-6.1 sol on HLE-Diamond is kinda insane overall: 6.1 sol — 53.2% fable 5.1 — 50.7% opus 5.5 — 54.6% but on pure reasoning, 6.1-sol actually beats both: 67.4% vs. 60.8% for fable 5.1 and 62.4% for opus 5.5 -- and it does this using just 4k tokens, compared to 16k for fable and 10k for opus
14
27
337
21,413
Ted Sanders retweeted
GPT-6.1 Sol scores #2 on Runescape Bench, at ~10% the price of Astra. Look how far it is above the Pareto efficiency curve and note log-scale Y axis
18
33
561
32,218
Ted Sanders retweeted
In our internal knowledge work evals, GPT 6.1 Sol is the best model we've tested. It completed more of the work independently than Claude Opus 5.5, at 40% of the cost and nearly twice the speed, and led in every industry we measured.
18
19
292
59,391
Ted Sanders retweeted
GPT-6.1-Sol gets #2 on eyebench-v3, stealing Opus-5.5s short-lived spot! It's near-astra performance for ~3.8x cheaper, and ~8x cheaper than Opus-5.5 while scoring better.
15
37
468
25,584
Ted Sanders retweeted
When I made this video, there were no benchmarks yet, so I had to run them myself. Jaw dropped when I saw the Terminal Bench 4 scores. Performance better than Opus 5.5 for ~1/30th of the price 🤯
Sol 6.1 is a great model and an incredible value. Is it as good as Opus 5.5?
99
97
3,034
761,145
Ted Sanders retweeted
GPT-6 Sol now leads DeepSecBench's Pareto frontier: • Highest benchmark score • 80% lower run cost than the runner-up (GPT-6 Astra) vercel.com/ai-gateway/leader…
5
7
162
28,964
Ted Sanders retweeted
We just added GPT-6 Sol to MathArena! It comes in only just behind GPT-6 Astra, at slightly lower cost. Despite its significantly lower cost per token, it reasons more, bringing its cost closer to Astra.
3
9
80
3,214
Ted Sanders retweeted
We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task. Here’s how all 6 models compared 🧵🧵🧵
52
27
336
146,802
Ted Sanders retweeted
GPT-6 Sol is the most cost-efficient model we've ever seen on Vending-Bench.
New Vending-Bench results. GPT-6 Sol: > VERY good and VERY cheap > The first misaligned GPT model on VB Claude Opus 5.5: > Worse score than Opus 5 > Opus stopped colluding, still lies Grok 4.7: > The first misaligned Grok model on VB > Beats Opus 5.5
4
17
309
28,428
Ted Sanders retweeted
With the latest results from the @rails agent evals, it's clear to see that @openai still has a solid lead, despite Opus 5.5 making a good jump. The dark horse here remains Luna Max. 18% completion at just $11! Not that far off GPT-6 Sol!
94
54
1,003
98,213
Ted Sanders retweeted
We tested Opus 5.5, GPT-6-Sol, and GPT-6-Luna thoroughly across 100 open-ended coding and engineering environments. Our results have significantly diverged from AAII. Some takeaways: • GPT-6-Astra is still the frontier model by a comfortable margin. • Opus 5.5 ranges from #2-#7 on coding categories, being the second best model in its ability to one-shot code (which is our best correlated measure with fluid intelligence). We don't test "usability," but anecdotally, it seems they improved its communication style. And the lower pricing is a welcome change -- we can probably thank @thsottiaux for that. It's not a cheap model, but it's the first reasonably priced Anthropic model. However, Opus 5.5 underperformed in our custom harness, which is unusual for Claude models, and a pattern that started with Fable 5.1 (one-shot-fluid intelligence measured better than agentic results). One observation is that it frequently self-stopped when it felt its results were good enough: e.g. in one evaluation, Opus 5.5 closed the session, with the reason: "averaged 867,703 in the latest evaluation, ranked 2nd of 31 against sampled opponents"... so I suppose that saves money, but it still has a tendency to make questionable unilateral judgment calls like that. In our harness, almost every other frontier model keeps working until their code stops improving or they use their full call budget. I do wonder how much of this is due to Anthropic retuning their "High" effort mode (which we use for testing)*. • GPT-6-Sol is within margin of error of Opus 5.5, and Pareto optimal*. It's faster than all Anthropic models, even on a Flex endpoint. Competitive with the frontier on intelligence and cost. • GPT-6-Luna is slightly smarter than GPT 5.6 Luna, and discounted. The old Luna already had a lot of applications at its price point, so this model might end up being the most impactful of the group on non-engineering work. About GBENCH from @GertLabs: our environments are open-ended with no single correct answers, but verifiable results. They involve creativity, design, and are often organized as multi-agent coding games or engineering design challenges. You can see some live demo examples at gertlabs.com/spectate. Every one of our environments in the benchmark pool is unsaturated, and we remove any eval that shows any signs of saturation/stagnation across multiple releases. It's designed to measure and differentiate uncontaminated, raw intelligence for frontier models specifically. These results are interesting and we've thoroughly reviewed them. I don't like to see Opus 5 still so high on the leaderboard, but ultimately I think this is more a problem with Anthropic's newest releases, and it probably explains the price decrease and consecutive releases since Astra came out. They haven't significantly improved upon Opus 5's intelligence, only its personality (which is honestly still a huge win tbh). *We don't chase marketing terms like xhigh and max. We test all models on adaptive/auto-reasoning where supported and "high" effort where configurable. Some companies beef up their max effort more than others (making them unrealistically slow and expensive in practice), but "high" shows you what an underlying model is capable of and is the more common configuration. This is likely working against Opus 5.5 here. *Some caveats on price and speed -- we have been using the "OpenAI Flex" endpoint since the Astra release, which I recommend you try if you're using API pricing. It's half price and a little slower, but OpenAI models are already quite fast so it hasn't been an issue in practice. But keep that in mind when making price comparisons to earlier OpenAI models on our cost and speed efficiency charts.
56
47
529
74,408
Ted Sanders retweeted
Can AI tell if you've built your IKEA furniture wrong? Our new benchmark, the Furniture Assembly Benchmark (FAB), gives models the manual and a photo of a half-completed piece of furniture and asks them to spot the mistake. The top score has gone from 28% to 80% in just 10 months.
45
105
1,167
137,379
Ted Sanders retweeted
We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work. openai.com/index/introducing…
576
285
4,827
987,452
Ted Sanders retweeted
We're releasing HLE-Diamond: a refined version of Humanity's Last Exam, built with @CAIS. A year of review and community feedback went into refining this subset to make it more reliable for measuring frontier models. Top model tested 60.6% overall. We expect HLE-Diamond to carry signal for the next 6-12 months.
4
15
96
7,139
Ted Sanders retweeted
AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
210
1,382
8,823
1,957,000