Exciting news: Claude Sonnet 5.5 (Max) by @AnthropicAI has debuted at #3 in the Agent Arena with +12.5% net improvement! This release is a 8.1 percentage-point increase over Claude Sonnet 5 (High), which ranks #13 with +4.4% net improvement. By category, Claude Sonnet 5.5 secured the #1 spot in Chat (+15.6%) above both Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%). This performance comes with a higher cost: Claude Sonnet 5.5 (Max) has a median cost of $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58. @AnthropicAI models now hold all three top positions in Agent Arena. Congrats to the team!
Exciting update: Claude Sonnet 5.5 with xHigh reasoning has landed in the Code Arena: WebDev. With 1786 pts, its ranked #3! At a blended $8/M tokens, Claude Sonnet 5.5 remains on the Pareto frontier with xHigh reasoning. This release is just 2 pts from GPT-6 Astra in the #2 spot with 1788 pts, for 80% of the price. By domain, Claude Sonnet 5.5 (xHigh) landed: - #2 in Gaming, Reference-Based Design, and Brand & Marketing - #3 in Simulations - #4 in Content Creation Tools and Consumer Product - #6 in Data & Analytics Congrats again to @AnthropicAI on this release!

Oct 2, 2026 · 7:31 PM UTC

24
19
311
37,637
Dive into the Code Arena: WebDev leaderboard at arena.ai/leaderboard/code/we…
1
5
3,239
Sort replies: Relevant Recent Liked
Replying to @arena @AnthropicAI
ngl i think its super cool seeing a whole family as either 123 or 456 lmfaooo nice
168
Replying to @arena @AnthropicAI
Ive been using sonnet 5.5 daily and its amazing, nice to see it ranked
34
Replying to @arena @AnthropicAI
i would guess astra 6.1 is ~ Fable.1 lvl
171
Replying to @arena @AnthropicAI
Anthropic is dominating the game
28
Replying to @arena @AnthropicAI
Agent Arena第三——Anthropic把'编程'这块招牌焊死了。榜单年年变,但有一条不变:写代码这件事,Claude系就没输过。OpenAI的下一张牌得加把劲了。
56
Replying to @arena @AnthropicAI
Makes sense. I use Sonnet 5.5 as worker under Opus right now and it’s been sooo good
233
Replying to @arena @AnthropicAI
@Grok que mide esto exactamente
1
156
Replying to @arena @AnthropicAI
i'd enter the arena, but my whole strategy is stealing shiny objects
192
Replying to @arena @AnthropicAI
包前三很爽,但 73% 成本溢价,把 Sonnet 推成了 Opus 价位。
53
Replying to @arena @AnthropicAI
the performance jump is reallly cool.. but that 73% higher cost can't be ignored!! the real question is how much that extra performance is worth per task😬
107
Replying to @arena @AnthropicAI
A Sonnet that costs 73% more per task than Opus is a fun inversion. Is that $2.74 coming from Max taking more steps per task, or from using more tokens on each step?
3
Replying to @arena @AnthropicAI
they're ranking agents. we ranked our chairs. tie for last.
58
Replying to @arena @AnthropicAI
Great point!
4
Replying to @arena @AnthropicAI
argon has been shit even before it publicly released
27
Replying to @arena @AnthropicAI
Claude Sonnet 5.5 Max gains significant ground at #3 in the Agent Arena benchmark
102