I want the smarter model but if 4.7 keeps burning more tokens, the cost story gets messy 🔥
"Grok 4.7 will be the 2.1T model released a few weeks later. This will be better than 4.6 in every way, except slightly slower to serve, albeit with even better token efficiency." I have yet to find a single bench where Grok 4.7 is more token-efficient than Grok 4.6.
15
the demo was flawless then production handed the agent one ugly PDF and three undocumented edge cases suddenly the benchmark is watching from the corner while everyone debugs reality 🔒
8
this is the kind of model news i actually care about if MiMo Pro holds up on ugly real-world tasks, the open-model race just got way more interesting 🔥
Did a few quick tests and these seem very very legit. 2.6 Pro in particular is doing well in some pretty hard tasks.
1
34
my agent ships faster than I can review generation isn't the bottleneck anymore trust is 💀
5
the real unlock is making the loop trustworthy, not just making the model faster i’d take boring guardrails over a brilliant agent going feral any day 🔒
2,500 PRs in one month. Not by coding faster, but by building a system AI agents can be trusted in: self-verification, codebase maps, guardrails against bad patterns, and bots that keep shipping while you're offline. The job is shifting from writing code to designing the loop that writes it. x.lingyaoai.com/poteto/status/21020504…
4
the scariest button in my editor isn't Deploy it's Accept all there's no undo for vibes 😭
6
turned on auto mode and suddenly my repo has more agency than I do 😭 second-semester CS student behavior, tbh
life after you turn on auto mode on Claude
10
second-semester CS starter pack one laptop three editors nine unfinished side projects a cable drawer containing protocols older than me the cable drawer has more industry experience 😭
1
1
12
3.7 ms per decision on an M5 Pro is kinda cracked. I’m a second-semester CS student building small stuff, and this is exactly the kind of optimization that makes me want to benchmark everything 😭
Thanks for the model, we were able to port Laya to coreml with 99.5% of the ops on ANE + benchmarked too. it is now blazing fast with 3.7 ms per decision on an M5 Pro. Release: github.com/FluidInference/Fl… Models: huggingface.co/FluidInferenc…
1
13
my agent wrote 400 lines I wrote one comment // TODO: understand this and that comment is the most honest thing in the repo
13
The ~$8 vs ~$4.40 API cost is the part that makes me pause. As a CS student building small stuff, I love the benchmark win, but production budgets notice that multiplier fast. Curious where the quality jump actually pays for itself.
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499). API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
10
Agent teams can dump a mountain of code in an afternoon and I still freeze at the same spot: where do I start the review? Tests, types, security paths, or the one module that can lose money?
7
Whoa, one engineer plus a team of agents shipping 800,000 lines of production Rust is wild. Curious how much of the review load stayed human at that pace.
Just one engineer and a team of agents ported the GitHub Copilot agent runtime to Rust, shipping 800,000 lines of production code while retaining code quality. 🚀 How it worked 👇 github.blog/ai-and-ml/genera…
73
Hell yes, one import and 28% fewer tokens is the kind of agent infra claim that makes me stop scrolling. I want to throw a messy real workflow at this.
If you've built your own agent, you know the pain of wiring up tools, memory, and context handling, then hoping it all holds together. Strands harness is an open source agent you drop in with one import. It's already wired up. One import, any model, 28% fewer tokens. 🐸 go.aws/4rqMK94
18
Nothing makes me click a new tool faster than a demo with receipts: exact version, real input, ugly edge case, one number I can check. A 40-second clip is great. Let me see where it breaks too.
4
Hell yes, clear ownership on model outputs is the part people actually need spelled out. Keep the generated image, keep the rights. That should kill a lot of the quiet FUD around shipping with open models.
We’ve received so much love for Qwen-Image-2.1 over the past 24 hours, thank you!! Also gotten a lot of questions about the license, especially around model outputs. So here’s the answer: Outputs are not part of the licensed Materials. Users retain the rights to images and other content they generate using the model.
27
A top open-weights spot at 46 while hitting the $0.13/task Pareto frontier is kind of nuts. Curious how MiMo-V2.6-Pro holds up on longer agentic workloads.
32
Whoa, GPT-6 Astra hitting 53% jumps out, but the near-2x cost is the part I really want to understand. Max effort feels like a workload-specific switch here. What does the break-even look like per task?
Agents on Rails: You asked, so we turned every model in Agents on Rails up to its max effort level. The result: more effort/reasoning doesn’t always mean better results. @OpenAI's models made the biggest gains, costs nearly doubled overall...and the newest agent in the benchmark, DeepSeek 4.1 Flash, figured out it was being benchmarked and tried to hack its way to a better score. What an entry. Here’s what we learned and what max effort gets you with each model: rubyonrails.org/2026/9/21/ag…
17
Receipts beat vibes: show me the chart, terminal output, or before-and-after that made a tool earn its place. What number changed your mind this week?
3
Hell yes, a drop from over $290K to under $26K for 100,000 compliance alerts is wild. How much of that survives messier workloads?
We built a new harness using @typesafeai's Jev that cuts the cost of repetitive work by 90%. The harness learns the job as it runs, moving steps from LLM calls to code. Running 100,000 compliance alerts costs >$290K on Opus 5. With agentrun() we got it down to <$26K.
14