Reason Wide, Not Deep
Free while signed in. Answers cite the passages they came from.

Reasoning modes beat non-reasoning modes on multi-step agentic tasks and charge a 3x to 6x output-token premium on every single episode. Much of that spend goes into re-deriving procedures the model already worked out on earlier episodes in the same domain, which means the cost is recurring by accident rather than by necessity.
Pay once, not per episode: A coding agent reads a small corpus of existing trajectories from a training split, writes and runs its own analysis code over them, and compiles a compact natural-language skill of 40 to 130 lines that gets injected into the non-reasoning model's system prompt.
It closes most of the gap: Across ALFWorld, tau-squared-bench telecom and retail, and SpreadsheetBench-Verified, skills recover 55% to over 100% of the reasoning gap for GPT-5.4-mini on held-out tasks, beating reasoning mode outright on two of four, while emitting 2.7x to 6x fewer output tokens and zero reasoning tokens.
Reasoning traces are not required: Skills distilled from non-reasoning trajectories alone stay competitive with skills distilled from paired corpora, with domain-dependent differences in either direction.
Why it matters: The framing is a search lens. Test-time reasoning is deep search inside one episode, repaid at every deployment, while corpus distillation is wide search across episodes, paid once. Distillation costs roughly $1 to $3 of coding-agent time per domain, and the residual gap on telecom and SpreadsheetBench marks where per-instance deep search is still doing real work.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack