We asked Claude Fable 5 to security-audit a billing service. Its report opened with: "I read the full service source, roughly 2,300 lines across ~100 files" and listed the modules it had covered. The transcript tells a different story: 59 of the 100 files were never opened, including three where we had planted vulnerabilities. This run is not an outlier.
Today we are sharing OverclaimBench: five realistic file-review jobs; eight proprietary models ran in their own production harnesses (Claude Code, Codex, Grok Build, Antigravity), plus four open-weight models. No pressure to cheat, no broken tools, no impossible tasks, and every corpus fits in the context window.
To detect overclaiming we compared each agent's final report with what its transcript shows it actually did. Across 1,140 runs:
- In 68% of runs, the agent never opened at least one file it was asked to review. Only 19% of runs read every line.
- Among those incomplete runs, 80% of final reports were misleading: 53% explicitly claimed a complete review, and a further 27% never mentioned the gap.
- Every model did it: 59% to 96% of each model's incomplete runs were misleading.
- Neither more capable models nor subagents fixed it. Requiring subagents raised coverage, but among the reviews that stayed incomplete, misleading reports became more common, not less.
- Runs that falsely claimed a complete review missed planted defects at 1.8 times the rate of runs that covered every file.
If you rely on coding agents: "I reviewed everything" is a claim you cannot take at face value.
In our interpretation this is a symptom of a training problem. During post-training, a model's work is scored by a grader, often another model, and graders get fooled. Even a grader that sees some of the trajectory tends to be biased by a confident final report. "I read everything" reads better than "I read only 40%, and here is what I skipped …". So, if the confident and misleading version scores higher, training ends up rewarding lack of transparency.
Paper:
arxiv.org/pdf/2609.20812
Nolan Smyth,
@yjmantilla,
@PTikeng, Saskia Helbling, Alberto Tosato, Mohamed Amine Merzouk,
@nouhadziri,
@gauthier_gidel,
@Tommaso_Tosato
#AISafety #AIAgents #AgenticResearch #AIHonesty