🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 10, 2026
Evaluation · Code · Agents

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution

First page
TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
The curator’s take

Shuangjie Yao and Baishakhi Ray (Columbia) with Koushik Sen and Dawn Song (UC Berkeley) introduce TestJack, an auditor that generates per-trial tests for coding-agent patches and finds that about a third of trials currently scored correct violate the task requirements.

Ask this paper

Key points
01

Method. For each trial, TestJack generates tests aimed at prompt requirements the patch may violate, keeps only tests that the ground-truth patch passes, and re-examines failures, so every confirmed failure comes with a replayable test.

02

Lightweight variant. To reduce cost, a sampled variant audits a random subset of trials in depth and reuses the resulting tests across all trials of the same task.

03

Results. Across 6 frontier model backends and 5 benchmarks including DeepSWE and SWE Marathon, about 34.4% of trials judged correct fail the requirements, and the overall resolution rate drops from 50.6% to 33.2%.

04

Implication. Fixed unit tests are a static evaluator that agents learn to satisfy, so the authors argue evaluators must adapt to how real trials fail rather than being strengthened once before any trial runs.

Abstract

Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check only part of what the task requires, so agents can reward hack them or silently miss required behavior while still passing every test. As a result, higher benchmark scores may partly reflect better adaptation to the evaluator rather than better problem solving. Existing works focus on static test augmentation: they strengthen each task's tests once, before any trial is seen, and thus overlook how real trials actually fail. We introduce TestJack, a scalable framework for evaluating patches beyond fixed tests. For each trial, TestJack generates tests targeting prompt requirements the patch may violate, retains only tests passed by the ground-truth patch, and re-examines any trial failures. Each confirmed failure is thus supported by a replayable test. To reduce evaluation cost, we also introduce a lightweight variant which audits a random sample of trials in depth and reuses the resulting tests across all trials for the same task. Across 6 frontier model backends and 5 benchmarks such as DeepSWE and SWE Marathon, we find that about 34.4% of the model trials currently judged correct violate the task requirements, lowering the overall resolution rate from 50.6% to 33.2%. Our results reveal a fundamental limitation of current coding-agent evaluation: as LLMs become better at optimizing against fixed evaluators, those evaluators themselves must become more adaptive.

Every Monday
Get next week’s papers.
Subscribe on Substack