Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499). API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40

Sep 21, 2026 · 10:57 PM UTC

85
98
954
4,212,239
Example 1 Task: PE firm's deck template requires a comparable-benchmarking section with a step-by-step valuation chain Grok 4.6 runs a brief analysis and lands exactly on the deal partner's shorthand estimate, whereas Grok 4.7 runs the valuation chain independently and flags the difference
2
43
25,313
Example 2 Task: PE firm's deck template requires an asset profile covering scale and financials Grok 4.6 uses only the latest year’s data with no revenue trend or currency discussion, whereas Grok 4.7 evaluates the full three-year history and concludes that growth is wiped out by currency depreciation, informing the buying decision
1
1
35
12,840
Sort replies: Relevant Recent Liked
Replying to @ArtificialAnlys
Awesome 👏
Grok 4.6 >>> Grok 4.5 Grok 4.7 >>> Grok 4.6 Grok 4.8 >>> Grok 4.7 Grok 4.9 >>> Grok 4.8 Grok 5.0 >>> Grok 4.9 Grok 5.1 >>> Grok 5.0 Grok 5.2 >>> Grok 5.1 Grok 5.3 >>> Grok 5.2 Grok 5.4 >>> Grok 5.3 Grok 5.5 >>> Grok 5.4 … … SOMEWHERE IN GROK 5.x IT HITS ESCAPE VELOCITY AND NO OTHER MODEL IN THE WORLD WILL BE ABLE TO MATCH IT
5
1,774
Replying to @ArtificialAnlys
Benchmarks have lost any credibility. This is sad
4
21
2,695
Replying to @ArtificialAnlys
damn. didn’t lie
Replying to @itslueul
Grok 4.7 should be roughly on par with Opus 5.0, not 5.1. Better in some ways, worse in others. We need to fix multimodal performance. Grok 4.8 will be a noticeable improvement. Grok 4.9 is probably Astra/Fable class. Grok 5 maybe better than anything. We shall see.
1
9
3,305
Replying to @ArtificialAnlys
~50% of opus cost and still right behind. wild
8
2,579
Replying to @ArtificialAnlys
The fact that Jensen said what he said is already a very good sign.
1
2
1,939
Replying to @ArtificialAnlys
Those error bars overlap across the whole top five.
2
1,314
Replying to @ArtificialAnlys
Im interested in Claude
2
819
Replying to @ArtificialAnlys
@grok I have Anthropic account for $100 per month with projects set up for work with instruction sets and skills, etc. are you telling me I should be using Grok 4.7 instead? Can you do the same type of thing as Claude AI? Projects or is it just even easier and better? Don't lie.
2
1
2,345
Replying to @ArtificialAnlys
Compared to Chatgpt, what are its advantages?
2
1
1,245
Replying to @ArtificialAnlys
4.7 was also faster in my testing today with rebuilding a massive powershell script. Was twice as fast at reviewing the code, finding bugs and recommending fixes. Then about 35% ish faster at writing the code.
3
1,555
Replying to @ArtificialAnlys
I am not feeling the fear politicians and others have tried to place into the AI race. Appears to me that all of these AI are trying to compete. So are the models thoroughly tested or not?
1
2
1,300
Replying to @ArtificialAnlys
Are they paying you to say this? Because the model is really bad...
1
1
370
Replying to @ArtificialAnlys
Gogogo @Grok! Save humanity from those evil other LLMs!
1
1
1,047
Replying to @ArtificialAnlys
I'm an Elon fanboy, but come on. The Grok I know is ChatGPT 3 or 4 level, it's really that bad on a conversational level still.
1
262
Replying to @ArtificialAnlys
We live in a society that treats people like absolute cattle. Anything that apeases the wall street whores is deemed ethical and as long as it makes money it does not matter who it harms. The ethics of #AI and yet we have interest pushing this globally. x.lingyaoai.com/unusual_whales/status/…?
Students who use AI almost every day score lower in science reading tests, per FT:
658
Replying to @ArtificialAnlys
@grok 这张图主要显示了Grok 4.7的最牛的优点是什么?
1
1,654
Replying to @ArtificialAnlys
分析质量 Elo 1698→1994 这跳挺狠……AA-Briefcase 这种长活测,比刷榜题更贴日常。成本砍一半也是真香点。
513
Replying to @ArtificialAnlys
Did y'all get paid to post this because Mimi v2.6 pro mogs this model on most fronts
1
3
528
Replying to @ArtificialAnlys
third on AA-Briefcase at about half Opus cost per task, with a clear jump from 4.6 on the Lite due-diligence set
2
381
Replying to @ArtificialAnlys
half the cost of opus on due diligence work. thats the line
2
554
Replying to @ArtificialAnlys
Honestly don’t even get a lot of the craze over financial modeling. They’re all not all that great at it, needs serious human oversight also. The main data providers have basically made this a non-issue.
2
439
Replying to @ArtificialAnlys
Grok 🔮✨
2
224
Replying to @ArtificialAnlys
if you are tired of nonsense benchmarks and do not trust them anymore check unbenchmark.com You can read and post reviews of 600+ AI models
1
698
Replying to @ArtificialAnlys
90pct of use cases are fine with gemini the model last
1
296
Replying to @ArtificialAnlys
YAY go Rocket man an Company for all the inspirating work and planetary leadership what vision ❤️🌈👍
2
193
Replying to @ArtificialAnlys
Awesome 👏
2
175
Replying to @ArtificialAnlys
The jump in Analytical Quality Elo is actually wild. If Grok 4.7 can deliver this level of reasoning at roughly half the cost of Opus 5, that’s a serious step forward. 🔥
1
318
Replying to @ArtificialAnlys
benchmaxed???
1
1
36
Replying to @ArtificialAnlys
The Three Pence Famine Diary Log 699b….21st September 2026 Back Home. Back in Liverpool. Our kitchen. Rain on the window, kettle clicking off, and me sat at the table with a bowl of cornflakes and the box propped up against the milk. Gobbolino was on the windowsill, tuxedo-black coat against the grey light, white chest and paws tucked neat, watching the rain like it owed him money. Dave was under the table, muscular white-and-brown bulk spread across my feet, chewing something he definitely shouldn't have. I turned the box around and read the back. Then I started doing sums on the corner of a newspaper. "Right, Gob. Listen to this." I drew a little lorry. "One artic. Full of cornflakes. London to Liverpool. That's about two hundred and thirty-five miles by road." Gobbolino's ear twitched. Dave stopped chewing. "That trailer takes twenty-two pallets on the bottom deck and twenty-two on top. Forty-four pallets. Five hundred boxes each. That's twenty-two thousand boxes of cornflakes on one wagon." I took a bite. "Now. That wagon does about eight miles to the gallon. Two hundred and thirty-five miles means about a hundred and thirty-four litres of diesel. At today's bulk price, call it a hundred and eighty-seven quid for the whole run." I wrote it out and circled it. "Divide that across the forty-four pallets. That's four pound twenty-five a pallet. Per box? Less than a penny. Fuel costs you less than a penny per box of cornflakes to get it from London to Liverpool." Dave's tail thumped once against the lino. "Now here's the bit they're setting up for us. They're telling us fuel is going to quadruple. So let's quadruple it. Diesel goes from a hundred and eighty-seven quid to seven hundred and forty-eight. Cost per pallet goes from four pound twenty-five to seventeen quid." I leaned over and showed Gobbolino the newspaper. "Extra cost per pallet? Twelve pounds seventy-five. Per box? Two point seven pence." I held up a single cornflake. "Two and a half pence. That's the real cost of the crisis. On a box you buy for two quid, the fuel shock is worth less than three pence." Gobbolino blinked slowly. He understood. He always does. "So what will they actually charge? They'll stand on the telly and tell you the shelves are empty, the ships aren't sailing, the wagons can't move. And they'll put that box up by sixty pence. Maybe more. Because the crisis is the best excuse they've ever been handed. They don't need to raise prices to cover costs. They need a story to hide behind while they raise them anyway." Dave let out a low rumble. He knows a bluff when he smells one. "Two point seven pence," I said, tapping the cornflake. "That's the truth of it. Everything above that is a decision. And the decision belongs to a boardroom, not a fuel tank." "Question everything," I said. "Especially a price rise with a story attached." The Three Travellers 🐾
1
444
Replying to @ArtificialAnlys
Separating analytical quality from presentation is useful. I'd add operator correction time to cost per task: minutes to verify the analysis, fix the deck and reach an accepted deliverable. The cheaper first pass can still be the more expensive outcome. // STATIC MONKEY
152
Replying to @ArtificialAnlys
@elonmusk , guys how are you finding or having this one way communication here because Mr Musk never replies back or available to read the other ppl’s answers/ thoughts/ concerns/ suggestions so its just one way traffic. I am not liking it.
8
Replying to @ArtificialAnlys
Benchmarks don’t matter to me nearly as much as real-world results. I want to see what an AI can actually produce through a complex collaboration—not what a chart claims it can do. The final product is the benchmark that matters. :)
480
Replying to @ArtificialAnlys
I would love to see a speed*intelligence benchmark to identify models that are both smart AND fast. from there it would be nice to compare cost vs speed*intelligence
103