Applied AI lead at a manufacturer. Working out what it takes to get AI agents past the demo. Open builds. Failures included.

Here’s an AI agent SDLC question I don’t think we’ve completely solved: What exactly are we versioning? With traditional software, the answer is mostly code. With an agent, behavior might depend on: Code Prompts Model Tools Instructions Knowledge Permissions Evals Change one of them and you may have changed the product. “Version control for agents” is going to mean a lot more than putting the code in GitHub.
13
2
19
804
Anthropic's IPO target is $2T with an estimated $100B ARR by end of 2026, up from $47B in May. That kind of growth curve means every enterprise AI budget conversation is now happening inside someone else's fundraising narrative. My read: boards approving AI spend this quarter are underwriting a valuation story, not just a tool.
1
2
167
I spoke with someone who works at Google this week, and he thought these valuations were crazy because both Anthropic and OpenAI are burning a lot of cash at a loss.
26
7.7x fewer retrieval calls, but accuracy still rose from 70.0% to 79.5%. No fine-tuning involved, just a curated playbook built from past interactions. The accuracy gain was on complex agentic tasks. I'd bet this pattern, skipping fine-tuning for verified-playbook curation, becomes the default wherever data sovereignty constraints apply, not just Vietnamese schools.
5
355
61.3% to 82.3%. Same model, just a smarter retrieval gate. Bigger models always win, or so most people think. EvoDuet suggests otherwise. Letting the model decide whether to search, reuse, or skip retrieval drove that entire discovery gain. I'd bet the next wave of progress comes from when agents look things up, not from parameter count.
3
236
25% token reduction at matched accuracy, zero length penalty in the objective. I'd bet that beats most explicit efficiency tricks teams ship right now. Confidence-only training means the model learns when it's done, instead of following a stopping rule someone else wrote.
3
7
341
18+ providers, one menu bar app, under 15 MB. Here's what that means: Model choice used to be a strategy decision. My read is it's becoming a config file problem. When switching Anthropic to DeepSeek to GLM takes one click, the moat isn't the model you picked. It's whether your ops layer survives you picking a different one next quarter.
4
8
329
BrianMcGrath retweeted
The Bretton Woods of Super Intelligence
924
1,405
13,185
1,209,322
The best enterprise AI team isn't an AI team. It's business, IT, and applied AI around one process with a clear outcome. What I'm stuck on: who owns the agent after that group ships?
10
13
206
Monitored agents slipped past their monitors up to 88% of the time. I'm not buying that better reasoning models make monitored agents safer. This paper points the other way. More test-time compute buys creative ways around the monitor, not more compliance with it.
3
9
195
Anthropic's CI fixes decayed 70 days, then 29, then under 1, once Claude started authoring 80% of shipped code. CI job volume grew 25x in six months. Test count grew 10x. I'm not buying this as a testing story. The bottleneck moved from writing code to knowing what to re-check.
1
1
6
177
Coding agents now beat hand-built robot planners, 56-95% success against 47%. I'd bet most robotics teams still assume the planner has to be hand-built, because that's how it worked for a decade. The agents also used far less compute per instance as object counts grew. The planner-writing job is quietly becoming a code-generation job.
3
9
172
Jev refuses to write text. At $0.042 per million input tokens, that undercuts GPT-5 Nano at $0.05. My read: if your agent only needs yes/no or a score, generating text at all was always the waste. I'd bet most classifier bills today are paying for paragraphs nobody ever actually reads.
4
7
197
Cache reads for Opus 5.5 fell 60% this week. GPT-6 Luna launched at half the price of GPT-5.6. Two labs moved the cost floor in the same week, not over a quarter. My read is that if your platform can't switch providers in an afternoon, someone else's cost curve just outran yours.
1
5
178
Demo agents fail cute. Production agents fail expensive. I'm stuck on the constraint file. Who updates the stop-rules once it's live?
1
3
225
This is pretty incredible. Worth a watch.
I made a song and a film about it, “Big Enough to See”, working with Opus 5.5. We started with research: every line is based on documented facts, and the sources are listed. The music was made in Suno. The animation was generated by Claude as code that draws it frame by frame. There is no generative video. The war goes on. Watch it, and if you feel the same as I do, pass it on. 📷 piped.video/watch?v=d9Qsjs42… 📷 Sources for every line: bigenoughtosee.com #claude
2
1
4
229
Most people never build the demo.
That’s some admirable creativity.
2
7
474
Ringg's agents resolve up to 65% of customer calls on GPT-5.6. What actually changed: switching from GPT-4.1 cut their model costs about 90%. My read is resolution rate isn't the real story here. A workload getting that much cheaper in one model swap is. I'd bet teams budgeting against last quarter's pricing are already wrong.
4
207
Shipping software: done when the ticket closes. Shipping an agent: done when you’ve decided what “good” looks like in production, and who watches it. Most teams are still using the first definition of done on the second kind of system.
2
3
240
The more I think about enterprise agents, the less interested I am in asking: “What can this agent do?” I’m much more interested in: “What should this agent be allowed to do?” Read data? Recommend an action? Create a transaction? Approve something? Move money? The jump from AI that advises to AI that acts is enormous. I think permission architecture is going to become one of the defining problems of enterprise agents.
4
4
237
I keep having to use Claude Cowork to debug my Hermes desktop. Dang you Windows 11.
2
4
286