Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1) 81.3% mean score (#1) APEX-Accounting: 15.4% Pass@1 (#1) 62.0% mean score (#1) On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree. Opus 5.5 on APEX domains: Management Consulting: 80.0% Pass@1 (#1) Corporate Law: 71.2% Pass@1 (#3) Investment Banking: 69.3% Pass@1 (#2) Accounting: 62.0% mean score (#1) Token use can indicate domains where the model is most effective. Consulting leads at 80% and uses the fewest tokens at 2.0M per attempt. Law has the highest partial credit (85.7% mean) but costs the most at 4.9M tokens. In Accounting, Opus 5.5 gets partial credit on most tasks but fully passes only 1 in 6. More tokens can increase capability on Opus 5.5. Max effort consumes 3.50M tokens per attempt. That’s 2.1x more than Opus 5 and 1.2x Fable 5.1. But list price fell to $4/$20 per M (Opus 5 was $5/$25), and cache reads are $0.20, so 2.1x the tokens is only about 1.7x the dollars. Effort level has a significant impact on benchmark scores and token usage. Medium effort: 52.3% on 756k tokens Max effort: 73.4% on 3.50M tokens Increasing effort to max gains 21 points for 4.6x the tokens on agentic work. On APEX-Accounting, the same jump buys 2.8 points for 3.4x the tokens. When Opus 5.5 solves a task, it solves it consistently. 142 of 239 tasks passed on all 4 runs. Failing runs used 3.1M tokens vs 1.9M for passing runs, and took twice as long. 13 runs hit a hard failure, all in Corporate Law. 8 of them burned 25M to 42M tokens before dying, showing that long-horizon legal work can still send the model into a loop. Congratulations to the @claudeai team. See full leaderboard: mercor.com/apex/apex-agents-…

Sep 22, 2026 · 5:04 PM UTC

9
4
62
21,587
Sort replies: Relevant Recent Liked
Replying to @mercor
👀
2
771
Replying to @mercor
This one is SOTA by a mile. Great job @AnthropicAI team!
1
249
Replying to @mercor
Agent能力涨得快,会计任务的稳定性还差一截
158
Replying to @mercor
Pass@1 on APEX-Agents is the clean number. The one I want is how much of that survives a harness change, since model and scaffold usually move together and the score won't tell you which one moved
77