🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 16, 2026
Code · Evaluation

Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead

First page
Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
The curator’s take

Fengshuo Liu (Imperial College London) and colleagues audit 254 published SWE-bench submissions without running any model and find that the top of the Verified leaderboard cannot be ordered from the published verdicts.

Ask this paper

Key points
01

Shared outcomes: The two leading Verified entries each resolve 396 of 500 instances; the top ten share 285 successes and 51 failures, so only 164 instances distinguish them, and frontier solution sets nest at 0.935 against a score-implied 0.774.

02

Scaffold effects: Holding the model fixed, scaffold choice moves scores by up to 29.8 points, larger than the 8.8-point spread of the top thirty; six of nine interaction tests stay significant after Holm correction.

03

No separable pairs: Exact paired McNemar tests separate none of the adjacent top-thirty pairs on Verified; a leader-based rule gives three descriptive tiers, or two after correction.

04

Recommendations: Report comparison-set-specific resolution, record the model-scaffold pair, publish tiers, and accept new instances by how many discordant outcomes they add; a five-step audit protocol is released.

Abstract

Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit whether the published verdicts support this reading, using 254 SWE-bench submissions across four splits without running models. On Verified, the leading two entries each resolve 396 of 500 instances. The top ten share 285 successes and 51 failures, leaving 164 instances that distinguish their outcomes. Frontier solution sets have median nesting 0.935 against a score-implied baseline of 0.774, indicating strongly shared successes. Scores also depend on the evaluated model-scaffold pair: observed within-model scaffold ranges reach 29.8 percentage points, compared with the 8.8-point spread of the top thirty. Six of nine cell-mean interaction tests remain significant after Holm correction, although this observational design does not identify causal scaffold effects. Exact paired McNemar tests separate none of the 29 adjacent Verified top-thirty pairs at alpha=0.05, while the larger Test split separates 14 of 23. A stated leader-based rule yields three descriptive tiers, or two after Holm correction; non-rejection does not establish equivalence. We release the partition and a five-step audit protocol that profiles shared outcomes, tests paired differences, reports grouping sensitivity, and estimates the instance budget needed for resolution. The results motivate reporting comparison-set-specific resolution and model-scaffold provenance instead of interpreting small aggregate gaps as established rank differences.

Every Monday
Get next week’s papers.
Subscribe on Substack