This 30-minute AI Engineer session by Nearform tech lead Alfonso Graziano shows how a coding agent can improve another agent through measured, reversible experiments.
He starts with a golden dataset and scorer, gives the coding agent a job file with the objective, repo, metrics, constraints, and relevant files, then tests one hypothesis per branch.
03:35 - define expected outputs and tool calls
13:31 - establish a baseline, test a hypothesis, keep the gain or roll back
15:29 - preserve each run in a branch, report.md, and shared memory
19:44 - turn user traces into failure clusters
26:09 - add every confirmed failure to regression tests
28:30 - connect specs, quality gates, context, and observability
In one production case, the eval score moved from 67% to 86% in about 10 iterations. The agent found edge cases, improved the system prompt and tool descriptions, and fixed tool logic.
Watch the session, then use the guide below to build the task contract, project map, typed tools, durable handoff, evaluator rubric, permission policy, and controller.