In our 72-run test, Sonnet 5.5 at max spent 23 cents of output for the same 40/40 high got for 2 cents. That waste sits inside every AI revenue line. We now block max. What still earns Opus surprised us. Is that priced in?
Every Claude Code sub-agent we ran was Opus 5.5. Then Sonnet 5.5 scored 40/40 for $0.02.
We never compromise on quality, so Opus built everything. Yesterday's test changed that: 72 runs, hidden tests, every effort level.
→ Sonnet 5.5 at high: 40/40, 2.0k tokens a task
→ at max: the same 40/40, 23k tokens
→ $0.02 of output at high, $0.23 at max
So now a Jev classifier reads each sub-task and picks the builder:
→ Sonnet 5.5 high for bounded work
→ Sonnet 5.5 xhigh for multi-file work
→ Opus 5.5 xhigh for money, auth, sends and live state
0.44 s and about $0.00004 a pick.
The twist: Sonnet is never allowed to run at max. A hook refuses it before the agent starts.
We open-sourced the whole setup: the agents, the effort levels, the router, the Jev picker and the guard.
Which of your sub-agents still run on max?
Repost this now, because most teams pay for max on work high already passes.
Follow for more measured Claude setups. Code in the reply.