Dogfooding the CLIs and MCP servers we're building has shown how tenacious agents are at working around problems, and how easy it is not to notice.
The first time I ran into this, we were just getting started on the CLI and creating a page under a parent was broken. The agent figured out the separate move operation still worked, so it created the page first and moved it under the parent afterwards. I wouldn’t have spotted the bug from the finished page. I only saw it because I was reading the actual tool calls it was making.
Another time, I wondered why my Confluence pages looked better than other people’s. My agent had dug through an internal repo for implementation details missing from our skill. We added those once we realized. Other users couldn’t access that repo, so they’d been missing instructions I didn’t even know my agent needed.
With a UI, I’m the one dealing with the friction. Agents won’t stop until they solve the problem. So they try another operation, find extra docs, then save it to memory for next time.
Daily use still matters, but it will lie to you if you only look at the final artifact. Just goes to show even more why isolated evals are so critical.
Anyone else finding their agent is sneakily working around bugs they didn’t know they had?