🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Evaluation · Code

Failure as a Process

Free while signed in. Answers cite the passages they came from.

First page
Failure as a Process
The curator’s take

When a coding agent fails a task, the final pass or fail label hides when the run actually went wrong. This large-scale study treats failure as a timeline and annotates over 63,000 execution steps to see how coding-agent runs break down.

Key points
01

Failure has three timestamps: Each trajectory is marked with the decisive error, the point where the error becomes irreversible, and the first observable failure, exposing a fix window and an observability lag that pass or fail labels erase.

02

Built on real trajectories: The team collected 3,843 runs from seven frontier models across three scaffolds on Terminal-Bench, filtered to 1,794 valid trajectories, and annotated them with high inter-rater agreement.

03

Mostly epistemic, mostly early: About 57.9% of failures come from misusing available information rather than a capability gap, with false premises the single largest trigger at 30.7%, and errors typically start early and stay hidden until recovery is impossible.

04

Why it matters: Naming the onset, lock-in, and observation points gives teams a vocabulary to intervene before an agent run is unrecoverable, instead of only noticing at the end.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack