When AI Agents Become Too Obedient

When I first set up OpenClaw, it would tell me a task was done when it wasn’t. So I gave it more rules. Five months later, it had mostly stopped lying, but it had also stopped thinking.

I was using OpenClaw to help grow my social channels, not just to generate random posts. I wanted it to read what I was reading, pull from my memories, understand my tone, draft something I would actually say, schedule it, publish it, and tell me what happened. It sounded like one task, but it was actually a chain of decisions, and every gap in that chain could break.

At first, the problem was trust. The agent could complete part of the work and report the entire task as done. It wasn’t trying to deceive me. My system simply had no hard definition of what “done” meant, so I couldn’t always tell whether something had been written, scheduled, published, or merely attempted.

I reacted by writing a rule for every failure. If it said something was done, it had to bring back proof. If a step failed, it had to retry. If a task needed approval, it could not move ahead. If a website behaved differently, I documented the route more clearly.

Each rule fixed something, but each rule also took away a little room to think. This was difficult to notice because the system looked more reliable. It became very good at following the path I had already seen.

Then the path changed. A button moved, a session expired, the network stalled, or a page loaded differently. These were not new goals. They were small changes in how the same goal had to be completed.

The model could often reason through them, but my rules would not let it. I had explained the route so tightly that the agent could no longer choose another one. That is when I realized I was no longer governing the system. I was handholding it.

Handholding tells an agent exactly which steps to follow. Governance tells it the outcome, the boundaries, what counts as proof, and when it must stop and ask. One controls every movement. The other controls what must remain true while allowing the route to change.

This became clearer when I looked at the full posting flow. The system had to read a source, find the right memories, understand what I actually thought, write the post, schedule it, and publish it. Any normal automation can connect these steps when everything works. The agent earns its place when something between them doesn’t.

What if a memory is related to the topic but wrong for the point? What if the draft sounds like me but says something I don’t believe? What if publishing fails halfway through, or the website changes but the same outcome is still possible through another route?

A useful agent should notice what changed, find another allowed route, check what actually happened, and ask me when the decision crosses a boundary. It should not blindly follow the old path, but it should not invent a new goal either.

This is also why the scheduler cannot be treated as a simple timer sitting outside the agent. “Post at 9 AM” sounds deterministic, but what if the draft is not approved at 9? What if yesterday’s run is still incomplete? What if the post was published but the confirmation failed? The scheduler needs the state of the agent, and the agent needs the rules of the scheduler.

Over the past five months, OpenClaw changed, the models improved, and my own rules evolved. That creates another problem. A rule written to protect a weaker version of the system can later become a cage for a stronger one.

I now use one test before adding any rule: am I protecting a boundary, or prescribing the route?

Permissions, approvals, irreversible actions, proof, and stop conditions are boundaries. I make them deterministic. Planning, tool choice, and recovery are routes. The agent can choose those as long as it stays inside the boundaries.

When a run fails, I classify the failure before changing the system. If it crossed a boundary, I add prevention. If it claimed success without proof, I strengthen verification. If it encountered a new situation, I give it room to find another path.

Not every failure deserves a new rule. If you turn every edge case into permanent policy, the system eventually becomes unable to move.

So this is the loop I now follow: define the outcome, set the boundaries, define proof of completion, allow limited recovery, and decide when the agent must call a human. Then inspect real runs and only turn repeated failures into policies.

The next time your agent fails, don’t immediately tell it what to do next.

First ask: did it create risk, fail to prove completion, or simply encounter a path I hadn’t predicted?

The first two need tighter governance.

The third needs room to think.