Prompt-Induced Waste
Free while signed in. Answers cite the passages they came from.

Two prompts can request the same code change and produce the same correct patch while causing a coding agent to perform radically different kinds and amounts of work. This preregistered study measures that effect across 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real harnesses.
Wording changes where effort goes: Prompt phrasing redirects effort into different work. Asking for multiple approaches inflates reasoning by 2.4x to 7.4x across all six open models and produces roughly three elaborated but discarded solution branches, still yielding one implemented solution and no success gain.
A second pathway runs through tools: Maximum certainty wording propagates into extra test runs, tool calls, turns, latency, and context growth. Runs with high redundant verification cost 18x the clean-run median, execute 2.5x more tool calls, and take 3x longer, again with no success gradient.
Harness design amplifies both: Cost per successful task swings by 5x to 30x in this setting depending on the harness, and the findings survive a frozen holdout, paraphrase tests, a Kimi-K3 replication, and a first-party Claude Sonnet 5 study.
Why it matters: Bounded-efficiency wording that specifies scope, acceptance criteria, and a stop condition preserves diagnosis and final validation while coming out neutral or better on all six holdout models, so most agent spend is decided before the model reasons at all and both levers are cheap to change.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack