Coding-agent benchmarks should grade repair, not first-pass success
Coding-agent benchmarks should grade repair, not first-pass success. A binary score for whether the first submitted patch passes tests rewards lucky completion and hides the capability teams depend on after a plausible attempt fails. Production coding agents must inspect evidence, localize the fault, revise the patch, preserve correct work, and avoid introducing a second regression. The repair loop is not a retry around the task; it is the task.
The benchmark landscape already contains the pieces. SWE-bench uses repository issues, containerized environments, and resolution tests. HumanEvalFix starts from buggy functions. API-Bank includes a plan-and-debug tier, while state-based benchmarks compare the resulting database or application state with the expected one. These designs recognize that tool use is not merely selecting an API or emitting code. The missing emphasis is a score that distinguishes a system that repairs from one that repeatedly samples until a patch happens to pass.
I would construct benchmark episodes with a known failing state and staged evidence. The agent receives an issue, repository, and test harness, then produces an initial change. Some episodes should expose a failed resolution test, a conflicting hidden test, a dependency-boundary mistake, or an intentionally misleading local symptom. The next phase grades whether the agent uses that evidence to update its diagnosis, changes the relevant files, and reaches a correct final repository state without discarding unrelated valid edits.
That requires more than a terminal pass bit. I want separate measures for localization, diagnosis change, regression avoidance, number and cost of repair cycles, preserved progress, and final state. A repaired patch that touches the right dependency boundary after one informative failure is different from ten blind rewrites. Test-execution reward remains useful, but it should sit inside a trajectory score that attributes improvement to evidence-guided correction rather than repeated lottery tickets.
The harness must also prevent benchmark gaming. Tests used for feedback should be separated from held-out acceptance tests, and the agent should not be able to rewrite the evaluator or weaken assertions. Each instance needs a clean container, controlled network access, and an attributable tool and model configuration. When scaffolding, parallel subagents, or extra compute are allowed, the benchmark should report them as part of the system. Otherwise the leaderboard conflates model capability with an undisclosed repair apparatus.
Repair-oriented grading teaches better engineering habits as well as measuring them. A learner—or an agent trained from benchmark trajectories—should see failure as structured evidence: which invariant broke, where the first divergence appeared, and which hypothesis the new result rejects. That supports the broader argument that agent QA should test recovery paths. It rewards making the next attempt more informed, not merely making another attempt.
There is one precise concession: first-pass success still matters for small, deterministic changes because each repair cycle adds latency, compute, and reviewer attention. I would keep it as a component of efficiency, not the definition of capability. A benchmark that ignores first-pass quality would excuse waste; a benchmark that ends at the first pass cannot measure resilience.
The acceptance question should be: after the first plausible patch failed, did the agent become a better engineer? Did it read the failure correctly, preserve what was sound, correct the causal defect, and leave the repository in a verified state? Coding work in production is dominated by that loop. Benchmarks should stop awarding the whole grade at the moment before the most revealing part begins.