New blog: Is human taste overrated in harness engineering?
An agent harness is the system around a model. It decides which tools the model can call, what evidence it sees, and what it remembers between steps. Most harnesses are still designed by human taste. Run the agent, read where it failed, change the harness, try again.
That loop is increasingly being automated, with a strong model rewriting the harness itself. So we asked how far that can go. We ran Qwen on three drug design tasks from SMDD-Bench and compared a manually redesigned harness against one produced by automated harness search with Claude in the loop.
No single approach won everywhere. The manual harness pulled clearly ahead when the fix was changing what the agent sees and remembers. Automated search did better once that was in place and the remaining problem was how to search the chemistry itself.
The hard part was rarely implementing the fix. It was diagnosing what kind of failure we were looking at, and that is where human taste still mattered.
Work with
@KevinH1119568 ,
@aviral_kumar2 and
@niloofar_mire
Read more 👇