Model or Harness
Free while signed in. Answers cite the passages they came from.

Agent evaluations mostly report system-level outcomes, so a failed run leaves the repair unassigned. The same visible failure might call for model post-training, harness engineering, environment redesign, or benchmark repair, and outcome labels cannot separate those cases.
Every failure gets an edge: The taxonomy organizes 41 failure modes by assigning each one to an edge between two components (model, harness, user, tools, memory, environment) plus a fault side naming where the repair belongs.
The schema is actionable by construction: Model-side failures identify post-training targets, harness-side failures point at scaffolding and tool-integration fixes, and environment or grader failures expose evaluation conditions that need redesign.
It survives automation: Across four frontier models, the strongest judge reaches Cohen's kappa of 0.76 against human category labels, so the labeling can run continuously over production traces instead of once per postmortem.
Why it matters: Harness engineering became the main lever for agent builders this year while teams still lacked a shared vocabulary for where a harness bug ends and a model bug begins. This supplies that vocabulary, and it applies across coding assistants, long-horizon personal assistants, and multi-agent systems.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack