Running an IT agency with AI agents, 14h a day. I share the real stories: what ships, what breaks. Built stepkeep, an SOP maker for Chrome ↓

Berlin
I run an IT agency with AI agents. The thing that ate our time: writing SOPs by hand. So I built stepkeep. Click through a task once, get a step-by-step guide with annotated screenshots. No account. No upload. No subscription. stepkeep.app
3
10
336
google says argon leads 12 of 18 of its own benchmarks, but terminal-bench 4.0 is 57.4% vs 66.4% for opus 5.5. agent work lives on terminal tasks, so for me the logo matters less than which model wins the task I actually run
3
42
best open-weight on cursorbench is a solid claim, and same price as 5.2 with no long context surcharge helps. i'd plug it into opencode and compare against sonnet 5.5 on real tasks
GLM 5.3 and GLM 5.3 Flash are now available in Cursor! GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0.
30
opus 5.5 cache reads dropped to $0.20 per million from $0.50 on opus 5. agents that re-read the same context all day pay that price most, so I think it matters more for running costs than the headline token price does
1
16
finding a bug through an agent and disclosing it privately first is how it should go. cursor hasn't confirmed it and there's no CVE yet, so i'd wait before calling it serious
We gave GLM‑5.3 a complex reverse-engineering task. It found a potentially serious vulnerability in Cursor. We disclosed it privately. Appreciate Cursor team is working closely with us on a fix, and we’ll share the more details once users are protected.
2
30
companies selling RL environments to AI labs, as seen from an agency dev
11
gemini 4 argon has a 1M token output limit vs 64K on earlier gemini, and my call is that quality past the first 100K is where it gets exposed. the benchmarks measure long context in, not long output, so a full repo refactor in one run is the real test
3
57
i'd still re-test everything after a point update, 6.1 sol being this different proves it. but these are their own cfo evals, so i'd run my own tasks before switching off opus 5.5
we found opus 5.5 to be 2X as efficient as gpt-6-sol in our internal cfo.ai evals so i was not expecting 6.1-sol to improve on that much but to see it 3X the efficiency of opus 5.5 and 2.5X against even sonnet 5.5 is actually 🤯 this is an insane point release
2
37
glm 5.3 uses the same 743B base as 5.2, and post-training alone took exploitbench from 24.4% to 54.4%. it's z.ai's own number, but it fits what I keep seeing: point updates are usually big changes. I'll test it in opencode next to opus 5.5
2
47
AI loves tests so much it keeps pinning the smallest stupid detail and test suite takes more and more time. It runs e2e even for a freaking font size change.
3
19
Tuxedo Winnie the Pooh, pricing page edition. Same if-statement, fancier label. Got Jev? Easier than you think
1
4
26
77.9 vs 74.2 on DeepSWE is google's own number, so i'll wait for independent runs. still, $2 in and $10 out with 1M output is cheap for agents. now put it on a subscription for enterprises
Very excited to announce Gemini 4 Argon, our new Frontier AI model, and a huge step forward across key capabilities. We’re focused on rolling it out responsibly starting with government and trusted cyber defenders through our Fairwind Program today, before wider availability soon
2
23
When fable 5.5 finally comes out it will change what we expect from our agents. The bar is constantly being pushed up. I loved glm 5.3, but i cant go back after i worked with Opus 5.5 and sonnet 5.5. Im afraid i wont be able to work with Opus once fable 5.5 drops
2
5
86
opus 5.5 leads epoch's index by 0.84 points and the 95% intervals overlap almost completely. that's a tie with astra, not a win. sonnet 5.5 at 165.2 is only 2 points back, which is why it's my default for most agent work
4
46
the cost per task gap is real, sol is $0.72 vs $5.98 for opus on the AA index. but opus is still 58 vs 52 on the same index and 65 vs 55 on terminal bench, so cheap isn't the same as better
we found opus 5.5 to be 2X as efficient as gpt-6-sol in our internal cfo.ai evals so i was not expecting 6.1-sol to improve on that much but to see it 3X the efficiency of opus 5.5 and 2.5X against even sonnet 5.5 is actually 🤯 this is an insane point release
3
74
4 cents per million input and free output is cheap enough to put in every agent loop. the accuracy numbers are perplexity's own though, so i'd run it on my own cases first
We’re open sourcing a state of the art multimodal Decision Model, pplx-decider-27b, and are offering it in a new Decisions API at 4 cents per million input tokens and free output tokens. We intend to bring down the price even further over the coming days. Enjoy!
2
5
74
44 for $6.56 vs 45 for $33 is the whole result. one test with planted bugs, but if it holds, sol is the one to run all day and astra only for the hard ones
So I tested GPT-6.1 Sol on real work: 2 repos, 105 planted bugs. Find and fix what you can. Unlike GPT-6 Sol, which was just a nerfed GPT-5.6 Terra, this one is real. The results: - GPT-6 Astra (max): 45 for $33 - GPT-6.1 Sol (max): 44 for $6.56 - GPT-5.6 Sol (max): 43.5 for $95.35 - Opus 5.5 (max): 41.7 for $58.53 - GPT-6 Sol (max): 29.3 for $9.33 With a model like this, the 50% cut to the $200 plan doesn't matter. n=1. More runs and results at more effort levels dropping in this thread over the next few hours 🧵
1
4
41
the report is impressions only, no clicks, and AI overviews and AI mode are lumped into one number. so it tells you that you show up, not that it brings anyone in
Google is now literally telling businesses specifically how to get traffic from AI Search. Yes, directly within Google Search Console. It's probably one of the most important marketing updates of 2026. Let’s go through it. By the way, you can see whether your business is appearing across Google AI, ChatGPT, Claude, Perplexity and Grok here (it's free): rightcited.com/ Recently Google launched a new Generative AI performance report inside Search Console. At first, it looked like a very limited rollout, but now more and more site owners are seeing it. John Mueller also confirmed that Google is rolling it out incrementally and reviewing feedback along the way. So if you do not see it yet, give it a few days, but the direction is pretty clear. AI Search visibility is becoming something businesses can actually measure, and this report shows how often URLs from your site appear in Google’s generative AI features. Yes, that includes AI Overviews and AI Mode. It can also show which pages appeared, which countries the impressions came from, which devices people were using and how performance changed over time. That matters a lot, because until now, a lot of AI Search tracking has been messy. It has featured a lot of screenshots, manual prompt testing, third party tools, guessing from referral traffic, watching ChatGPT, Perplexity, Claude, Grok and Google manually to see whether your brand appears, etc. Now Google is giving site owners a direct view into whether their pages are appearing inside Google’s own AI features. There is one big limitation, though: this is impression data as opposed to clicks data. Google is showing whether your site appeared inside generative AI features, but it is not yet showing how many people clicked from those AI features to your site. This is where SEO Stuff (seo-stuff.com/gold-plan-pack…) can help. A business can be visible in AI Search and still see fewer clicks than it used to get from traditional search. At the same time, a business can also be completely missing from AI Search and not realize competitors are being shown instead. If your AI impressions are high but clicks are weak, the issue may be the way Google is satisfying the search before the user visits your site. If your AI impressions are low, the issue may be that Google does not understand, trust or retrieve your pages for the questions customers are asking. If only your informational pages appear, but your product, service, pricing, comparison and case study pages do not, that tells you something. If one country shows AI visibility and another does not, that tells you something. If a few pages get most of the AI impressions, that tells you something. This is why the report is such a big deal, and it gives businesses a starting point so you can finally ask better questions. Which pages are showing up in AI Overviews? Which pages are showing up in AI Mode? Which countries are seeing the brand? Which devices are seeing the brand? Which parts of the buying journey are missing? Which content is actually being used by Google’s AI systems? That last question is the one most businesses should care about, and it is the one SEO Stuff solves for brands on a daily basis: seo-stuff.com/premium-conten… A customer can ask a long, complicated question and Google can break that question into several related searches behind the scenes. That means the same business may need useful content around problems, use cases, comparisons, pricing, alternatives, industries, objections, case studies, product details and implementation. One generic service page probably will not cover all of that. A brand with useful content across the full buying journey has more chances to be discovered. Google’s new report will make that easier to see. And this also lines up with what Google just said in its “Good SEO is good GEO” piece. (See my post from yesterday if you missed it.) Google said its generative AI features are rooted in the same core ranking and quality systems as traditional Search. AI Mode and AI Overviews retrieve up-to-date content from Google’s existing search index. So the businesses that want AI visibility still need the same boring foundation: crawlable content, useful pages, clear product information, real expertise, entity clarity, authority, third party validation, etc. That is where SEO Stuff (seo-stuff.com) comes in. The done-for-you package combines 10 AI-search-optimized articles with three DR50+ authority placements: seo-stuff.com/gold-plan-pack… The content helps your business cover the questions customers ask before buying, including problems, use cases, comparisons, pricing, alternatives, industries, objections, case studies, product details and implementation. The authority placements help your business show up across trusted sources that search and AI systems use to understand categories. Google is starting to show businesses whether they are appearing in AI Search. Once something becomes a report, it becomes something clients, founders, CMOs and competitors start paying attention to. If you want to see whether your business is already being cited, understood and recommended across Google AI, ChatGPT, Claude, Perplexity and Grok, check here: seo-stuff.com/free-audit
2
32
fine, but the slow week is the real cost. if sol is back to normal speed now, i'd like to see the tok/s numbers again before i trust it for agents
Global reset landing tomorrow 10am PST for all paid ChatGPT accounts. Apologies for the slow start with GPT-6.1 Sol, it's now back to running at expected speeds after the massive load spike in the first two days.
2
34
one user measured gpt-6.1 sol xhigh at about 20 tok/s vs 54 for gpt-6 sol. it's one post so take it carefully, but if it holds, sol is a batch job model. cheap and slow doesn't work for interactive agents
4
52