Organizing human intelligence to power the AI economy.

San Francisco
While it will AI will greatly benefit society, it will also displace hundreds of millions of jobs. We need better benchmarks to help politicians and enterprises measure model performance and prepare for the displacement.
New research: we hired 12 CPAs to complete four realistic accounting tasks, then had AI models from the past few years complete the same work. As of May 2025, accountants outperformed AI. Now they lose out to today's frontier models, which are near perfect on these tasks. ๐Ÿงต
2
4
34
5,787
Gemini 4 Argon is the new #1 on APEX-SWE Integration. Pass@1 scores for coding tasks: Integration: 72.5% (#1) Observability: 26.8% (#25) Overall: 49.6% (#13) It is the best Gemini model on APEX-SWE, +7.6 pts over Gemini 3.7 Flash. ๐—ง๐˜„๐—ผ ๐˜ƒ๐—ฒ๐—ฟ๐˜† ๐—ฑ๐—ถ๐—ณ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐˜ ๐—ฟ๐—ฒ๐˜€๐˜‚๐—น๐˜๐˜€ Integration tasks ask the agent to build end-to-end systems across services. Argon leads this domain, +3.2 pts over Sonnet 5.5 (69.3%). Observability tasks ask the agent to debug production failures from logs and telemetry. Argon scores 26.8%, 43 pts behind the leader, Opus 5.5 (69.8%). ๐—–๐—ผ๐—ป๐˜€๐—ถ๐˜€๐˜๐—ฒ๐—ป๐—ฐ๐˜† We ran every task 4 times. On integration, it passed 72 of 100 tasks on all 4 runs. Observability, it only passed 17 of 100 tasks on all 4 runs. ๐—ง๐—ผ๐—ธ๐—ฒ๐—ป๐˜€ Argon reads a lot more on Observability tasks: Integration: 1.3M tokens per attempt Observability: 7.5M tokens per attempt On Integration, failing runs used more than twice the tokens of passing runs (1.7M vs 0.7M median). On Observability, passing and failing runs used about the same (7.1M vs 7.8M median). More reading did not lead to more passes. ๐—™๐—ฎ๐—บ๐—ถ๐—น๐˜† ๐—ฝ๐—ฟ๐—ผ๐—ด๐—ฟ๐—ฒ๐˜€๐˜€ Gemini on APEX-SWE, Pass@1: Gemini 3.1 Pro: 33.9% Gemini 3.5 Flash: 36.1% Gemini 3.6 Flash: 39.4% Gemini 3.7 Flash: 42.0% Gemini 3.8 Flash: 36.3% Gemini 4 Argon: 49.6% Argon is +15.7 pts over Gemini 3.1 Pro. Congrats to @Google and @GoogleDeepMind.
13
7
158
8,766
APEX turns 1 today. Last year, we launched the AI Productivity Index (APEX) to answer whether frontier AI models can do professional work. Since then, weโ€™ve created new benchmarks while model capabilities progress more rapidly than even the boldest predictions. APEX remains the industry standard for evaluating frontier AI on economically valuable work. Hereโ€™s how APEX has evolved over the past year. ๐—ข๐—ฐ๐˜ ๐Ÿฎ๐Ÿฌ๐Ÿฎ๐Ÿฑ: ๐—”๐—ฃ๐—˜๐—ซ Mercor introduces APEX, our first AI benchmark for testing investment banking, law, consulting, and medicine. ๐—๐—ฎ๐—ป ๐Ÿฎ๐Ÿฌ๐Ÿฎ๐Ÿฒ: ๐—”๐—ฃ๐—˜๐—ซ-๐—”๐—ด๐—ฒ๐—ป๐˜๐˜€ Built with partners @box and @harvey, APEX-Agents evaluates AI agents on long-horizon tasks in investment banking, consulting, and corporate law. ๐— ๐—ฎ๐—ฟ ๐Ÿฎ๐Ÿฌ๐Ÿฎ๐Ÿฒ: ๐—”๐—ฃ๐—˜๐—ซ-๐—ฆ๐—ช๐—˜ Co-developed with @cognition, APEX-SWE assesses real production engineering across integration and observability. ๐—๐˜‚๐—น ๐Ÿฎ๐Ÿฌ๐Ÿฎ๐Ÿฒ: ๐—”๐—ฃ๐—˜๐—ซ-๐—”๐—ฐ๐—ฐ๐—ผ๐˜‚๐—ป๐˜๐—ถ๐—ป๐—ด Built with @tryramp and @RampLabs, APEX-Accounting measures whether models can close the books. ๐—ฆ๐—ฒ๐—ฝ ๐Ÿฎ๐Ÿฌ๐Ÿฎ๐Ÿฒ: ๐—”๐—ฃ๐—˜๐—ซ-๐—”๐—ด๐—ฒ๐—ป๐˜๐˜€ ๐Ÿญ.๐Ÿญ Our first major update to APEX-Agents includes newly audited tasks, an improved judge, and enhancements to ensure leaderboard accuracy. Thank you to all our partners who built these benchmarks with us. And, to the Mercor experts, professional bankers, lawyers, consultants, doctors, engineers, and accountants, whose judgement and expertise ensure these benchmarks are realistic. The rate of AI progress is only increasing. Mercor is committed to maintaining and extending our APEX family of benchmarks as the industry standard for informing decisions about the AI frontier and its ability to do professional work. See full leaderboards: mercor.com/apex/
3
9
49
3,572
Mercor retweeted
New research: we hired 12 CPAs to complete four realistic accounting tasks, then had AI models from the past few years complete the same work. As of May 2025, accountants outperformed AI. Now they lose out to today's frontier models, which are near perfect on these tasks. ๐Ÿงต
41
143
1,381
246,991
We compared 12 licensed accountants against frontier AI models on realistic accounting tasks. The models were faster and more accurate than every accountant in our study. 18 months ago, the best models scored below the average accountant's ~37%. Today, models ace the same tasks. This doesnโ€™t mean accountants are replaceable. But it does suggest the job will change significantly, even if model progress stalled today. Read our full study: mercor.com/blog/human-baseliโ€ฆ
5
40
189
57,944
Gemini 4 Argon is the new #1 on APEX-Agents. 82.2% Pass@1 (#1) 87.9% Mean score (#1) It is the first model to exceed 80% Pass@1. It is +6.7 pts over the prior #1, Sonnet 5.5, and +14.4 pts over the best prior model from DeepMind, Gemini 3.7 Flash (67.8%). #๐Ÿญ ๐—ถ๐—ป ๐—ฎ๐—น๐—น ๐—”๐—ฃ๐—˜๐—ซ-๐—”๐—ด๐—ฒ๐—ป๐˜๐˜€ ๐—ฑ๐—ผ๐—บ๐—ฎ๐—ถ๐—ป๐˜€ Google says Argon delivers frontier performance in "enterprise knowledge work like legal and finance." On APEX-Agents, Gemini 4 Argon ranks first in every professional domain. Management consulting: 90.3% (#1) Investment banking: 80.9% (#1) Corporate law: 75.3% (#1) Consulting appears to be Argon's strongest domain, scoring 10.3 pts ahead of Opus 5.5. It is also the first model to score above 90% on any APEX-Agents leaderboard. ๐—ง๐—ผ๐—ธ๐—ฒ๐—ป ๐˜‚๐˜€๐—ฎ๐—ด๐—ฒ ๐—ฏ๐˜† ๐—”๐—ฃ๐—˜๐—ซ-๐—”๐—ด๐—ฒ๐—ป๐˜๐˜€ ๐—ฑ๐—ผ๐—บ๐—ฎ๐—ถ๐—ป Argon used 2.6M tokens per attempt on average. The median attempt used 1.75M. Investment banking: 3.5M per attempt Corporate law: 2.7M Management consulting: 1.7M Almost all Gemini 4 Argonโ€™s token usage comes from input. It used 2.61M input tokens per attempt compared with only 28.5K output tokens per attempt. That is about 90 input tokens for every output token. The agent reads a lot of files before composing a short answer. Congrats to @Google and @GoogleDeepMind. See full leaderboard: mercor.com/apex/apex-agents-โ€ฆ
8
39
485
15,384
Voice agents need to work for everyone, whatever language we speak, however our names are spelled, wherever we live, and whatever specialized words we use. Two new papers from @SierraPlatform's ฯ„-voice team help put a number on that: ฯ„-Elicitation: Benchmarking for multi-turn entity extraction in voice agents ฯ„-Multilingual: Extends ฯ„-voice to a multilingual settings Proud that Mercor's Victor Barres was a key contributor. Congrats to the whole team! Read the full papers. arxiv.org/abs/2609.13602v1 arxiv.org/abs/2609.35820v1
Introducing ฯ„-Entity, a benchmark for how voice agents collect, verify, and correct the details callers give them. In ฯ„-Voice, capturing names and other identifying information was a recurring bottleneck: get a detail wrong, and authentication fails before the agent can help. ฯ„-Entity breaks that exchange down. Can an agent collect one exact value, recognize when it needs to check what it heard, and repair a mistake? We test 200 tasks across ten entity types, including names, phone numbers, addresses, and emails. Each call asks for a single field, with a cooperative caller who spells accurately when asked. Most calls last less than two minutes. And even here, reliability is far from solved. When agents choose their own verification strategy, only 14โ€“41% of tasks succeed across all tested conditions. A brute-force approachโ€”prescribing spelling, read-back, and confirmationโ€”raises that to 37โ€“54%, but makes calls longer. The challenge is finding the balance: knowing when to trust a capture and when to spend more effort checking it. Across systems, agents recognize hard and unfamiliar entities and verify them more carefully. But they do not increase verification for the caller voices they struggle with most, and their response to noise is inconsistent. And verification often fails to produce a repair: just 24โ€“37% of wrong captures that agents verify end up correct. Agents need to get better at both deciding when to check and using the answer to fix the record. Weโ€™re introducing ฯ„-Entity alongside ฯ„-i18n, and plan to extend these entity-collection tasks across languages with the upcoming codebase release. Together, they help us measure whether voice agents can both handle the details accurately and deliver a conversation that feels native. Thanks to Victor Barres at @mercor for collaborating on this. Papers: ฯ„-Entity arxiv.org/abs/2609.13602v1, ฯ„-i18n arxiv.org/abs/2609.35820v1 Codebase and multilingual extension coming soon!
1
26
3,544
GPT-6.1 Sol scores 60.0% Pass@1 and 73.2% mean on APEX-Agents (#11), and 46.9% on APEX-SWE (#18). @OpenAI says it nearly matches GPT-6 Astra on professional work "at one-fifth of Astra's standard input and output token prices." That is consistent with what weโ€™re seeing on APEX, within 4.7 points of Astra on Agents and 3.1 points on SWE. ๐—”๐—ฃ๐—˜๐—ซ-๐—”๐—ด๐—ฒ๐—ป๐˜๐˜€ GPT-6.1 Sol improves Pass@1 across all APEX-Agents domains compared with GPT-6 Sol. Corporate law: 63.4% (#11), up 7.2 Management consulting: 58.4% (#12), up 6.8 Investment banking: 58.1% (#10), up 2.9 That makes GPT-6.1 Sol OpenAI's #2 model on APEX-Agents, behind Astra and ahead of GPT-5.6 Terra (58.2%). It improves 5.7 points over GPT-6 Sol overall. On investment banking, GPT-6.1 Sol beats Astra (55.9%). It is now OpenAI's top model there. For corporate law, GPT-6.1 Sol trails Astra by 10 points on Pass@1, but its mean score is 85.3%. Only 4% of law runs score zero. It does most of the work and misses a rubric item or two. ๐—ง๐—ผ๐—ธ๐—ฒ๐—ป๐˜€ ๐—ฎ๐—ป๐—ฑ ๐—ฐ๐—ผ๐—ป๐˜€๐—ถ๐˜€๐˜๐—ฒ๐—ป๐—ฐ๐˜† 0.50M tokens per attempt on APEX-Agents. 1.29M per attempt on APEX-SWE. On law, failing runs use 2.4x the tokens of passing runs. GPT-6.1 Sol passed 49% of APEX-Agents tasks on all 4 runs. On APEX-SWE, 39%. 70 Agents tasks were never solved in any run. ๐—”๐—ฃ๐—˜๐—ซ-๐—ฆ๐—ช๐—˜ GPT-6.1 Sol improves slightly over the prior generation with 46.9% (#18), up 1.9 from GPT-6 Sol. The gain is all in Observability where it scored 32.3% (#19), increasing by 4.5. Integration was mostly flat at 61.5%, scoring slightly below GPT-6 Sol (62.3%) and GPT-6 Astra (62.0%). Congrats @OpenAI on the launch. See full leaderboards: mercor.com/apex/
6
4
53
3,315
Claude Sonnet 5.5 debuts at #1 on APEX-Agents and #2 on APEX-SWE. APEX-Agents: 75.5% Pass@1 (#1) 21 pt gain from from 54.5% for Sonnet 5 APEX-SWE: 66.4% Pass@1 (#2) 20 pt gain from 46.4% for Sonnet 5 Anthropic states Sonnet 5.5 at Max effort performs comparably to Opus 5.5. On APEX-Agents, it scores 2.0 pts higher (75.5% vs. 73.5%). APEX-Agents shows big improvement over Sonnet 5, where two domains gain more than 20 points. Investment banking: 77.2% Pass@1 (#1), up from 54.1% (+23.1 pts) Management consulting: 79.7% (#2), up from 47.8% (+31.9 pts) Corporate law: 69.6% (#5), up from 61.6% (+8.0 pts) In investment banking, Sonnet 5.5 takes over #1 from Gemini 3.7 Flash (71.3%). Consulting is the biggest jump over Sonnet 5, gaining +31.9 pts over the prior generation. On APEX-SWE, Sonnet 5.5 is 1.2 pts behind Opus 5.5 (67.6%). Integration: 69.3% (#1), up from 60.3% (+9.0 pts) Observability: 63.5% (tied #2), up from 32.5% (+31.0 pts) Sonnet 5 solved about 1 in 3 production debugging tasks, and now Sonnet 5.5 solves nearly 2 in 3. Congrats to @AnthropicAI on the launch. See full leaderboard: mercor.com/apex/
3
3
45
41,428
Claude Opus 5.5 is the new #1 on APEX-SWE. Overall, Opus 5.5 scores 67.6% Pass@1, a +3.9 point gain over the previous leader, Opus 5, at 63.7%. APEX-SWE measures observability, debugging from production telemetry, and integration, if a model can build an end-to-end system. Compared with Opus 5: Observability Opus 5.5: 69.8% (#1) Opus 5: 63.5% (#2) It gains +6.3 points. Opus 5.5 leads the board by 6.3 over Opus 5 and 10.8 over Fable 5.1. Integration Opus 5.5: 65.5% (#4) Opus 5: 64.0% (#9) Opus 5.5 improved at reading telemetry and finding faults, but building systems from scratch moved less. Congrats to @AnthropicAI. APEX-SWE leaderboard: mercor.com/apex/apex-swe-leaโ€ฆ
11
6
39
3,357
Mercor is committing $5M to a new AI Capabilities Fund. In conversations with researchers and enterprises, one challenge that comes up repeatedly is that many AI evals do not reflect how models are used. But to provide useful signal for eval and training, benchmarks need to capture actual workflows in representative environments. This will only become more difficult as models are used for more high-value, complex tasks. The Capabilities Fund supports researchers working on frontier eval challenges across a range of domains and data shapes, including: - Reward calibration and aligning grades with expert preferences - Creating high-quality realistic environments - Long-horizon workflows - Ambiguous requests and under-specified outcomes - Navigating social, temporal, and business context We will fund researcher time, API credits, and travel. We will also cover the cost to work with our network of 5M+ experts, as well as free use of our evals and analysis platform. Grants are available to independent researchers, non-profits, and academics. To submit an Expression of Interest, go to: mercor.com/careers/?ashby_jiโ€ฆ
7
5
41
8,943
OpenAI GPTโ€‘6 Sol and Luna are on the APEX leaderboards. GPT-6 Sol scores 54.3% Pass@1 (#18) on APEX-Agents. That is +2.9 pts over GPT-5.6 Sol. GPT-6 Luna scores 44.3% Pass@1 (#29), +1.3 pts over GPT-5.6 Luna. On APEX-SWE, Luna scores 38.8% Pass@1, +6.7 pts. OpenAI reduced API prices for Sol and Luna by 50% compared with their GPTโ€‘5.6 promotional pricing. GPT-6 Sol improves in finance and consulting domains: Investment banking: 55.2% Pass@1 (#13), +5.5 pts over GPT-5.6 Sol Management consulting: 51.6% (#16), +9.0 pts GPT-6 Sol is only 0.7 pts behind GPT-6 Astra in investment banking. These gains are not consistent across all domains. Compared with GPT-5.6, Solโ€™s corporate law score drops 5.7 pts to 56.2% (#27). Luna drops 2.8 pts to 55.6%. GPT-6 Luna improves most on coding work: On APEX-SWE, GPT-6 Luna scores 38.8% Pass@1, gaining +6.7 pts over GPT-5.6. Integration: 59.3%, +13.3 pts Observability: 18.3%, no change Sol scores 45.0% Pass@1 (#18), about level with GPT-5.6 Sol (45.8%). Integration: 62.3% (#13), level with GPT-6 Astra (62.0%) Observability: 27.8% (#19), down 3.7 pts For both models, debugging with production telemetry is still the hardest part. Congrats to the @OpenAI team on the launch. See full leaderboard: mercor.com/apex/apex-agents-โ€ฆ
5
1
31
2,382
Mercor has now partnered with > 300 enterprises to help them monetize their data, paying companies up to $10M per enterprise. Across these companies, weโ€™ve encountered 90 different systems: -55% contained Slack -53% GitHub -53% Google Drive -50% Gmail -40% Figma -35% Jira A company with 200 employees runs on many of the same systems as an enterprise with 5,000+. For AI to do economically valuable work, it needs to understand how work actually happens across these systems โ€” not just within any one of them.
24
24
368
2,138,299
APEX is now on the Specialized Intelligence Index, @FireworksAI_HQ's new leaderboard for real-work benchmarks by industry. We built APEX to test frontier modelsโ€™ ability to do professional work. Weโ€™re proud to partner with Fireworks to make these benchmarks more widely available. APEX General Practitioner (MD) APEX-Agents Corporate Law APEX-SWE Find out which models lead each category. See the full results on SII: fireworks.ai/specialized-intโ€ฆ
Today we're launching the Specialized Intelligence Index (SII): one destination for real-work benchmarks across industries, built by the teams that use them every day. Hear from Fireworks co-founder @the_bunny_chen on the importance of specialized benchmarks:
4
25
3,436
Claude Opus 5.5 is the new #1 on APEX-Agents and APEX-Accounting. APEX-Agents: 73.5% Pass@1 (#1) 81.3% mean score (#1) APEX-Accounting: 15.4% Pass@1 (#1) 62.0% mean score (#1) On APEX-Agents, the new model gains +4.9 pp over Fable 5.1 (68.6%), the previous leader, and +7.6 pp over Opus 5 (65.8%). Anthropic says the biggest gains in this release are on long-running agentic tasks and knowledge work. That is what APEX-Agents measures, and the numbers agree. Opus 5.5 on APEX domains: Management Consulting: 80.0% Pass@1 (#1) Corporate Law: 71.2% Pass@1 (#3) Investment Banking: 69.3% Pass@1 (#2) Accounting: 62.0% mean score (#1) Token use can indicate domains where the model is most effective. Consulting leads at 80% and uses the fewest tokens at 2.0M per attempt. Law has the highest partial credit (85.7% mean) but costs the most at 4.9M tokens. In Accounting, Opus 5.5 gets partial credit on most tasks but fully passes only 1 in 6. More tokens can increase capability on Opus 5.5. Max effort consumes 3.50M tokens per attempt. Thatโ€™s 2.1x more than Opus 5 and 1.2x Fable 5.1. But list price fell to $4/$20 per M (Opus 5 was $5/$25), and cache reads are $0.20, so 2.1x the tokens is only about 1.7x the dollars. Effort level has a significant impact on benchmark scores and token usage. Medium effort: 52.3% on 756k tokens Max effort: 73.4% on 3.50M tokens Increasing effort to max gains 21 points for 4.6x the tokens on agentic work. On APEX-Accounting, the same jump buys 2.8 points for 3.4x the tokens. When Opus 5.5 solves a task, it solves it consistently. 142 of 239 tasks passed on all 4 runs. Failing runs used 3.1M tokens vs 1.9M for passing runs, and took twice as long. 13 runs hit a hard failure, all in Corporate Law. 8 of them burned 25M to 42M tokens before dying, showing that long-horizon legal work can still send the model into a loop. Congratulations to the @claudeai team. See full leaderboard: mercor.com/apex/apex-agents-โ€ฆ
9
4
62
21,614
Mercor retweeted
Donโ€™t miss newly announced Fireworks Forge speaker: @BrendanFoody, CEO and Co-founder of @Mercor. Join the people, teams, and companies building their own frontier on open models. Nov 3, San Francisco. Apply to attend: fireworks.ai/forge#apply
1
4
16
3,878
Grok 4.7 is now live on the APEX leaderboards. APEX-SWE: 53.6% Pass@1 (#5) APEX-Agents: 54.6% Pass@1 (#13) Compared to other models with similar cost and latency profiles, itโ€™s a strong model for agentic coding tasks. Congrats to the @SpaceXAI team.
2
2
33
14,629
Grok 4.7 is very token efficient. On APEX-SWE, it uses 46% fewer completion tokens per task than Grok 4.6 (23.8k vs 44.1k). Full leaderboard: mercor.com/apex/apex-swe-leaโ€ฆ
1
11
2,273
For long-horizon tasks on APEX-Agents, Grok 4.7 is much faster than its full-size sibling, Grok 4.6. Built for speed, Grok 4.7 runs each APEX-Agents task with 60% fewer completion tokens and finishes 2.5x faster than its full-size sibling, Grok 4.6. Median time per task falls to 3.1 minutes from 7.7.
1
2
294
Four different Flash models now rank within the top 15 on APEX-Agents. Gemini 3.7 Flash: 67.8% (#2) Gemini 3.8 Flash: 64.3% (#6) Grok 4.7: 54.6% (#13) GLM-5.3-Flash: 52.8% (#15) Grok 4.7 leads GLM-5.3-Flash by 1.8 points on 240 tasks. Full leaderboard: mercor.com/apex/apex-agents-โ€ฆ
1
2
233