Failure as a Process
Free while signed in. Answers cite the passages they came from.

When a coding agent fails a task, the final pass or fail label hides when the run actually went wrong. This large-scale study treats failure as a timeline and annotates over 63,000 execution steps to see how coding-agent runs break down.
Failure has three timestamps: Each trajectory is marked with the decisive error, the point where the error becomes irreversible, and the first observable failure, exposing a fix window and an observability lag that pass or fail labels erase.
Built on real trajectories: The team collected 3,843 runs from seven frontier models across three scaffolds on Terminal-Bench, filtered to 1,794 valid trajectories, and annotated them with high inter-rater agreement.
Mostly epistemic, mostly early: About 57.9% of failures come from misusing available information rather than a capability gap, with false premises the single largest trigger at 30.7%, and errors typically start early and stay hidden until recovery is impossible.
Why it matters: Naming the onset, lock-in, and observation points gives teams a vocabulary to intervene before an agent run is unrecoverable, instead of only noticing at the end.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack