Building with AI coding agents and posting what actually works. AI @AudaciousHQ. Ex-CTO @GoCodeoAI. IIT Kanpur

SF
every new session i re-explained the same port, the same constraint, the same fix we already ruled out. the agents were never the bottleneck. i was the memory between them. so i built coding brain.
Article

Coding Brain: One memory for every AI coding agent

I run about 40 projects. Most days I have coding sessions open in parallel across Claude Code and Cursor, sometimes Codex too. The agents are good. That was never the problem. The problem was that

21
1
27
9,480
openai dots are ai agents that learn your workflow and handle tasks without asking. sama says his gets better daily as it learns what he doesn't like doing. i think the change is agents remembering your patterns, not just answering questions.
OpenAI
35
models do change behavior with updates and people notice when their prompts stop working. that's not a conspiracy, it's version drift. the labs aren't nerfing on purpose but they're also not locking quality once you pay for a tier.
If you think models are actually getting nerfed by the big labs as some weird method to force you into newer models, you should legitimately consider deleting this app. Thinking in such conspiratorial ways is bad for your mental health, and X algo will gladly help you spiral.
1
32
claude's showing a watermill someone built with opus 5.5 by writing rendering code. same method as those viral video posts. i think more creative tools end up code-first than most people expect.
Lots of inventive things made with Claude these past few weeks. A few of our favorites: A working watermill, built with Opus 5.5.
43
asd-ste100 is a simplified english spec from aerospace with 900 approved words and strict rules. karpathy's trick: ask llms to explain in asd-ste100 format and outputs get way clearer. i think this beats most prompt engineering hacks i've seen.
1
51
ai companies demo booking flights and planning weddings to show off agents. but those aren't everyday tasks, and they're complex enough that failures look bad. i'd rather see demos of small things people do ten times a day.
there are so many everyday demos that could make ai, personal agents, & agents in general feel insanely powerful to normal ppl & yet ai companies keep choosing things like booking flights, planning travel, or god forbid planning a god damn wedding. these are not everyday problems. they’re relatively rare events & worse for a lot of ppl they’re actually part of the joys of life. the best demo should make someone feel pain they already are familiar with disappearing. almost every company gets this wrong. they optimize for spectacle instead of relief which is fine in some cases but it doesn’t last. one of my favorite historical examples is a product called head on. the ad showed someone with a headache, then showed them applying the product directly to their forehead while repeating: “head on. apply directly to the forehead.” beautifully simple. you instantly understood the product, the problem, & the value. ai companies should be doing the same thing. show me the annoying thing i deal with every single day then make it disappear. it’s not sexy, i guarantee it will work. when jobs demo’ed the iphone he picked universal simple problems that the iphone did better than anything else on the fucking planet. anyway, thanks for coming to my ted talk.
34
ai models now score near 100 percent on accounting tasks. junior accountants averaged 37 percent. claude costs 21 cents per answer vs ten dollars for humans. i think most entry-level work crosses this line in the next year.
42
llms can output diagrams, html pages, or full explainer videos instead of text. asking for a custom video on any topic actually works now. i think we'll see a lot more throwaway web apps and videos built just to explain one thing.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
59
tavus says griffin is the first ai to pass the video turing test. 48 percent of people who talked to it live thought it was a real person, earlier systems were under 3 percent. i think this is where video ai stops feeling like a demo.
Tavus
91
fable 5.1 runs coding agents with less supervision than the alternatives. bindu moved 100 percent of her workflows to it. i'd test it on debugging work first, that's where supervision costs really add up.
Moved back 100% of our coding agents to Fable 5.1 It’s heads and shoulders above the rest and requires the least supervision Still, it’s far from perfect! Can’t wait for Fable 5.5
68
claude code shipped mods today. you write typescript to change ui, behavior, or add features, install with /plugin. cursor forked vscode to make an ai editor. this lets you mod the ai editor without forking it.
ClaudeDevs
1
230
deepmind built watermarks for ai-generated proteins. dna labs can verify if a protein came from ai and which safeguards were applied. i think this becomes critical once model-generated biology hits production scale.
SynthID Bio is our new family of watermarking methods made for AI-generated biological designs. In a world first, we can now embed an imperceptible signature directly into protein sequences without affecting their biological function. 🧵
44
grok bots can hand off coding to cursor and manage the whole pr flow now. you tell the bot what to build in slack, it writes the code in cursor and opens the pr. this is agents actually shipping code, not just suggesting it.
36
openai tracked how small businesses deploy ai agents across sales, product, and finance. makes sense - a 10 person team can ship an agent in a week while big companies are still writing the policy on whether agents can read slack.
Small teams are taking on more with AI—from finding customers to building products and managing finances. Our new report explores how small businesses are putting AI agents to work. And through a new partnership with @ASBDC, we're bringing hands-on AI training and local guidance to help more owners get started. openai.com/index/helping-sma…
43
figma makes developers wait 8 months to approve mcp access, amazon blocks shopping agents, reddit and x won't let claude or chatgpt read their content. the same companies that built on open apis are now locking them down. we're doing the api wars all over again.
1
1
2
549
google released gemini 4 argon, a new frontier model built for coding, enterprise work, and cybersecurity. it hit number one in text arena today. rolling out to testers through a closed program, so real performance data should show up soon.
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
78
grok bot shipped team bots you can share with your team, plaid integration for finance questions, and voice calls that search your chat history while you talk. team sharing is the piece that makes it useful for actual work.
31
cursor added /visualize that builds charts right in your chat. ask about data and it draws the answer inline instead of dumping text. i'd use this over asking for plotting code every time.
Cursor can now build charts and diagrams right in the chat. Use /visualize to analyze data and see the answer inline. Available now in the Agents Window.
68
openai launched dots, always-on agents that run on their own cloud computer and work across 4,000+ apps. you set what a dot does alone, when it asks first, and what it never touches. i think those permission settings decide who trusts it with real work.
OpenAI
72
openai now sells speed as its own tier. ultrafast runs its models up to 8 times faster, about 300 tokens a second in codex, and costs more per token. i think coding agents make this worth paying for, since waiting on the model is the slowest part of the work now.
54
meta, openai, and x all launched personal ai assistants this month. muse, dots, and grokbot want to be the one you use every day. i think whichever one integrates with the most apps wins, not whoever built the best model.
Muse vs Dots vs Grokbot Zuck vs Altman vs Musk. The battle of personal agents for the masses. Who do you think wins?
52