Scaling Laws for Agent Harnesses
Free while signed in. Answers cite the passages they came from.

Most harness tuning treats every token and tool call as if volume is what counts. This paper shows that it mostly does not, and introduces Effective Feedback Compute (EFC), a trace-level scaling coordinate that credits feedback only when it is informative, valid, non-redundant, and retained for later decisions, then normalizes by task demand.
Raw budget barely predicts success: In controlled scaling, raw tokens and tool calls explain limited variation in outcomes, with R-squared of 0.33 and 0.42. The usual cost proxies are weak predictors of whether the agent actually succeeds.
Effective feedback nearly explains everything: Oracle-EFC normalized by task demand reaches an R-squared of 0.99. Once you measure feedback that is genuinely useful and retained, the scaling behavior becomes almost fully predictable.
Quality beats quantity at fixed budget: In matched-budget interventions, improving feedback quality raises success from 0.27 to 0.90 while raw cost and tool calls stay fixed. The win comes from better feedback, not more of it.
Why it matters: Harness scaling is governed less by how much compute you spend than by how efficiently raw budget converts into durable, task-sufficient feedback. That reframes harness engineering as a feedback-quality problem and gives a coordinate to optimize against.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack