🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 26, 2026
Agents · Evaluation · Data

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

First page
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
The curator’s take

Edward Lue Chee Lip, Ivan Bercovich and colleagues audit a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record to ask what it means when no agent solves a benchmark task.

Ask this paper

Key points
01

Corpus. 1,081 pull requests, 639 scored tasks, 28,801 trials and $105,933 of logged agent spend, with reference-solution runs, empty-solution controls, adversarial trials, trajectories and review records.

02

Ordered validity screen. For each of the 125 tasks with no honest pass, the screen checks whether the reference solution passes, whether infrastructure failures dominate, whether the verifier can be bypassed, and whether solvability is supported by evidence.

03

Result. Only 78 of the 125 tasks survive as certified-unsolved candidates. 14 have broken oracles, 8 are dominated by infrastructure failures, 4 are passable only through verifier bypasses, and 21 have no evidence that they are solvable.

04

Narrow label. Certified-unsolved means the authored route passed, infrastructure did not dominate, no strict bypass was seen and every evaluated agent failed. The authors state it does not prove intrinsic hardness or verifier completeness.

05

Recommendation. Frontier benchmarks should publish the evidence behind their all-fail tasks before those tasks are used as capability claims, since pass rate alone does not explain difficulty.

Abstract

Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be bypassed. In this paper, we study this issue using a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record with 1,081 pull requests, 639 scored tasks, 28,801 trials, and $105,933 in logged agent spend. We ask what an all-fail task actually certifies. For the 125 tasks with no honest pass, we combine task artifacts, reference-solution runs, empty-solution controls, adversarial trials, trajectories, telemetry, and review records, and apply an ordered validity screen. Only 78 of the 125 tasks survive as certified-unsolved candidates. The remaining tasks include 14 with broken oracles, 8 dominated by infrastructure failures, 4 that are only passable through verifier bypasses, and 21 whose solvability is not certified by the available evidence. Thus, lack of saturation and genuine difficulty are not the same thing. The certified-unsolved label is also narrow: it means that the authored route passed, infrastructure did not dominate, no strict bypass was observed, and all evaluated agents failed. It does not prove intrinsic hardness, verifier completeness, or failure at the intended capability. We further analyze rejected submissions and passing tasks to show that pass rate alone cannot explain why a task is difficult. Overall, our results suggest that frontier benchmarks should report the evidence behind their all-fail tasks before using them as capability claims.

Every Monday
Get next week’s papers.
Subscribe on Substack