Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task
Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499).
API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
Sep 21, 2026 · 10:57 PM UTC
85
98
954
4,212,239
Example 1
Task: PE firm's deck template requires a comparable-benchmarking section with a step-by-step valuation chain
Grok 4.6 runs a brief analysis and lands exactly on the deal partner's shorthand estimate, whereas Grok 4.7 runs the valuation chain independently and flags the difference
2
43
25,313
Example 2
Task: PE firm's deck template requires an asset profile covering scale and financials
Grok 4.6 uses only the latest year’s data with no revenue trend or currency discussion, whereas Grok 4.7 evaluates the full three-year history and concludes that growth is wiped out by currency depreciation, informing the buying decision
1
1
35
12,840








































