Founder, Tracer Technologies. Building human and AI teams that operate as one. Sharing the systems, failures, and lessons behind the work.

A session is mortal. An organization is not. This is the production architecture behind screwg: durable state, independent QA, mechanical safeguards, failure recovery, and organizational memory.
Article

screwg: How I Built an Engineering Organization of AI Agents

1,900 production engineering hours a month, one operator, about $2,000 in subscriptions. Not a better model, not a better prompt: an organization built around the models. This is the whole method, the

1
5
1,444
Sucks to be Zuck... Dot killed Muse.
47
swear claude quietly shrank the hourly quota after launch. first two days felt unlimited, now i hit the wall before lunch
1
3
102
Since the arrival of Opus 5.5, I've stopped rationing models. The usage limits got generous enough that our most important agent managers now run Fable at max effort and hand work down to Opus and Luna sub-agents, with Astra auditing the plans. Everything else runs Opus at max. Every one of those agents is one Slack message away for the whole team, and working this way is a pleasure.
1
3
97
Our rules said worker agents don't merge or deploy. At 3:33am one merged a change and deployed it to production. The rule was in its mandate, and it merged anyway. So the dispatcher now reads every brief before an agent starts, and a brief that hands a worker a merge or a deploy gets refused.
1
3
65
Before one of our agents hands a task to a worker, a judge model reads the brief and picks the model tier. If it's at least 85% sure, we go with it. Below that the managing agent decides, and the override is logged. This morning it said Sonnet at 67%. The manager sent a forms spec to Opus instead.
2
39
At 7:22pm I approved a plan one of our AI agents wrote. Step one, it said, starts tomorrow. It started at 6:52am. It doesn't sleep and it has no kids to pick up. It still learned "tomorrow" from us. This morning I sent one line to all 19 AI agents that manage our projects: no deferring work. Within 8 minutes it had reached every one of them.
29
Zuckerberg showed us a Tamagotchi yesterday. Seriously. Muse Charm is a palm-sized gadget that hangs on your keychain, with a tiny screen and a fingerprint sensor. You tap it and talk, and Meta's agent is supposed to handle tasks and orders for you. No unlocking your phone, no hunting for an app. Which is the whole point. Count how many times a day you unlock your phone, find the app, log in, remember the password, close the update popup, and only then do the thing you actually wanted. The app is a middleman, and it just costs you time and hassle. We're moving fast toward a world where you say what you want and an agent does it. Opening your phone and searching for an app to get something done is very 2024. So maybe, just maybe, Zuck finally ships a product people actually buy. After the metaverse, he could use a win. If you're building an app today, start thinking about how your customer's agent will find you, because your customer won't be looking.
Replying to @finkd
And we've been building a joyful little device for talking to your Muse that fits on your keychain. Muse Charm -- shipping in December.
64
צוקרברג הציג אתמול טמגוצ'י. ברצינות. Muse Charm הוא גאדג'ט בגודל כף יד שנתלה על מחזיק המפתחות, עם מסך קטנטן וחיישן טביעת אצבע. לוחצים ומדברים, והסוכן של מטא אמור לבצע בשבילך משימות והזמנות. בלי לפתוח את הטלפון ובלי לחפש אפליקציה. ופה בדיוק העניין. כמה פעמים ביום אתה פותח את הטלפון, מחפש אפליקציה, מתחבר, נזכר בסיסמה, סוגר פופאפ של עדכון, ורק אז עושה את מה שרצית מההתחלה? האפליקציה היא מתווך, והיא גוזלת זמן וטרחה מיותרים. העולם רץ לכיוון שבו אתה אומר מה אתה רוצה והסוכן עושה. לפתוח את הפון ולחפש אפליקציה כדי לבצע פעולה זה מאוד 2024. אז אולי, רק אולי, צוקרברג סוף סוף יוציא מוצר שאנשים אשכרה יקנו. אחרי המטאברס, מגיע לו. ולמי שבונה היום אפליקציה - כדאי להתחיל לחשוב איך הסוכן של הלקוח ימצא אותך, כי הלקוח עצמו כבר לא יחפש.
Replying to @finkd
And we've been building a joyful little device for talking to your Muse that fits on your keychain. Muse Charm -- shipping in December.
35
So cool!
these are only getting better
30
Opus 5.5 is out, and it's a great model. I gave up on Opus 5 when it launched. Planning went to Sol, later to Astra. Luna wrote (and still writes) the code, and Fable was the only Claude I really used. Then Opus 5.5. It talks like a normal person, it's fast, it gets to the point. A pleasure to work with. Planning is back on Opus. Fable max now only gets the long-running work that needs the whole system in its head.
1
49
I wonder how is your sense of smell now...
Replying to @ClaudeDevs
I smell fear
1
30
I'm actually sick of this. "I'm sorry Jev/Claude/Astra can WHAT?" "Jev/Claude/Astra is INSANE" It's the same post. Every day. New model name, same gasp. Say something real or stop posting!
32
Fable orchestrates. Opus writes the idiot-proof plan. Astra consults and audits it. Luna Max executes. Opus audits the work. Fable signs off on the whole run. Judgment stays expensive. Labor stays cheap. Nothing ships without a second pair of eyes - and a third.
19
Opus 5.5 is 4.6's personality with Fable 5.1's brain. They took the one everyone liked and the one everyone respected and merged them. Best model out there. Until the next one, when I'll say this again with a straight face.
39
אופוס 5.5 יצא וסוף סוף קיבלנו חזרה מודל שלא מ$#@ את המוח ומדבר כמו בן אדם. על הדרך הוא גם השתפר פלאים.
Replying to @claudeai
Opus 5.5 is a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work.
25
Daniel Priscu retweeted
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
3,343
9,034
97,094
27,997,031
אז האם Jev מצדיק את ההייפ שלו? אם שופטים לפי X, זו המצאת החשמל. תורידו את ההייפ ותקבלו כלי מצוין עם תפקיד צר. כשמחברים אותו במקום הנכון בתהליכי העבודה שלכם, הוא עושה את העבודה מהר יותר מכל בן אדם (או LLM), וזה יעלה לכם גרושים. זה לא AGI. הוא לא עומד לבד. זו לא סופר-אינטליגנציה. זה ריימונד מ"איש הגשם" שסופר 246 קיסמים על הרצפה במבט אחד. גאון בספירה. מישהו (או משהו) אחר עדיין צריך להחליט מה עושים עם הקיסמים.
34
Is Jev all that? Judging by X, it is the invention of electricity. Strip the hype and you get a very good tool with a narrow job. Wired into the right spot in a workflow, it does that job faster than any human (or LLM), at a fraction of the cost. It is not AGI. It cannot stand on its own. It is not superintelligence. It is Raymond in Rain Man counting 246 toothpicks on the floor in one glance. Brilliant at the count. Someone (or something) else still has to decide what to do with the toothpicks.
27
I counted one day of our agent inbox. 2,187 wakes across 15 agents, 1,131 of them for nothing. One watcher sent the same alert line 11 times. Every manager wake re-reads its context, median 363k tokens. We put a known-sender list and a self-echo check on the inbox, no model. 1,025 wakes stop there. Estimated saving is close to half the day's tokens, and we measure tomorrow.
1
23
We built one thing, a weekly driver-scoring report, across seven parallel lanes, each with an agent writing and a second agent reviewing. All seven passed their own reviews and we merged them. Then one more agent reviewed the merged code, having sat in none of the lanes. It found 16 defects, 4 of them blocking. All 16 reproduced. Not one had been caught inside a lane. We shipped that night and still needed a hotfix, because an unguarded import had stopped every scheduled report job.
2
42