One agent file cuts Claude Code usage by 25% on a feature I built with 72 agents over 64 hours. Computed call by call on the real trace.
Unlike the main convo, sub-agents only keep their prompt cache for 5 minutes, even on a subscription. If one waits longer than that on a test run or a build, its next call pays for its whole context again. In that build it happened 200 times: 31% of the total cost.
The worst was one agent polling a slow job every 10 minutes. Each check cost ~20x a normal call (chart in the reply).
The fix is a sub-agent with a 1-hour cache. Save this as ~/.claude/agents/long-cache.md:
---
name: long-cache
description: General-purpose agent whose prompt cache lasts one hour instead of five minutes. Use it for agentic work that will sit idle more than five minutes between its own turns (waiting on long builds, test suites, visual checks, or its own subagents), especially once its context grows large, since every expired cache rewrites the whole context. For short or continuously active tasks use general-purpose: its cache writes cost less.
experimental:
cacheTtl: 1h
---
Work as a general-purpose agent on the task you're given.
Then ask Claude to use long-cache for anything that runs your slow commands.
Why not give every sub-agent a 1-hour cache? Its writes cost 60% more, and 53 of my 72 agents never waited 5 minutes. For them it's pure waste.