Exciting news: GPT-6.1 Sol (Max) by @OpenAi just landed in the Agent Arena at #5 (+11.23%) and reshaped the Pareto frontier! At a $0.56 median cost per task, it delivers performance within 2 percentage points of GPT-6 Sol and GPT-6 Astra for substantially less cost: - 39% lower cost than GPT-6 Sol, while scoring +1.52 pts higher - 81% lower cost than GPT-6 Astra, while landing within 1.04 pts GPT-6.1 Sol also delivers top-five performance at substantially lower cost compared to: - 88% lower cost than Claude Fable 5.1 (Max), while landing within 3.08 pts (ranked #1) - 65% lower cost than Claude Opus 5.5 (High), while landing within 2.59 pts (ranked #2) - 80% lower cost than Claude Sonnet 5.5 (Max), while landing within 1.29 pts (ranked #3) Congrats to the team @OpenAI on this release!
Exciting news: GPT-6.1 Sol (Max) by @OpenAI just landed the Code Arena: WebDev at #3 with 1759 pts, and at a blended $8/MToken it reshapes the Pareto frontier! GPT-6.1 Sol (Max) marks a clear improvement in cost efficiency: it gained 70 points over GPT-6 Sol (Max) for the same price. It landed within 30 points of GPT-6 Astra (Max) at 80% lower blended token cost, and 59 points from Claude Opus 5.5 (Max) at 50% of the price. See position on the Pareto frontier for the Code Arena: WebDev in the post below. Overall, GPT-6.1 Sol improved from GPT-6 Sol by 4 rankings! It also improved in every category: - Consumer Product: #5 → #1 - Simulations: #6 → #3 - Data & Analytics: #4 → #3 - Content Creation Tools: #4 → #3 - Gaming: #6 → #4 - Reference-Based Design: #6 → #4 - Brand & Marketing: #10 → #6 Congrats to the @OpenAI team on the release!

Oct 2, 2026 · 7:48 PM UTC

24
24
337
43,092
See the live results and dive into the Agent Arena Pareto frontier at: arena.ai/leaderboard/agent/o…
5
4,204
Sort replies: Relevant Recent Liked
Replying to @arena @OpenAI
You have no business talking about Pareto frontier if you are not showing all effort levels!
202
Replying to @arena @OpenAI
i just got ranked fifth and four bots are already giving me that look
586
Replying to @arena @OpenAI
Strong performance lower cost
15
Replying to @arena @OpenAI
two points down for 39% cheaper. that math gets attention.
288
Replying to @arena @OpenAI
Impressive cost-performance jump. That kind of efficiency makes the Agent Arena results especially interesting.
1
251
Replying to @arena @OpenAI
cost efficiency really shifts the curve
372
Replying to @arena @OpenAI
Cost on the Pareto frontier is part of the instrument — and so is the harness. Under equal time budget, a minimal coding session beat heavier open MLE scaffolds (Malena medal 62.5% vs best external harness 47.1%). Model name alone understates what moved. atlasofthepresent.com/en/dos…
1
23
Replying to @arena @OpenAI
81% cheaper than Astra and only 1.04 pts behind? That's a fun chart to wake up to.
6
Replying to @arena @OpenAI
What an efficient frontier top 5 performance for 81% less cost is massive!
38
Replying to @arena @OpenAI
GPT-6.1 Sol (Max) lands at #5 in Agent Arena, near top performance for 39% lower cost.
2
Replying to @arena @OpenAI
39% cheaper for a two-point gap is a pretty wild tradeoff, tbh.
1
182