🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 6 – Sep 6, 2026
Reasoning

It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories

First page
It's the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories
The curator’s take

Yigit Utku Bulut supplies the two counterfactual controls that the breakthrough-moment and early-legible-fate readings of reasoning traces have been missing, and both readings largely fail to survive them.

Ask this paper

Key points
01

Restart-controlled truncation probe: compares continuing from a prefix against generating from scratch at matched total token budget, separating a solution that merely fits the continuation budget from a prefix that carries value fresh compute cannot buy.

02

Exactly 1 of 178 cells is prefix-limited: across 89 MATH problems and two small open models, almost nothing behaves like a genuine breakthrough prefix. What looks like insight is mostly compute compression.

03

Continuing still beats restarting wherever the matched budget falls inside the restart grid, 9 of 9, which is a real practical finding even as it deflates the interpretive one.

04

Difficulty is the confound in early-window probes: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations, inside the published probe range, and a matched reconstruction of a published positive is indistinguishable from chance within problem, 0.496 at t=4.

05

The methodological demand: high pooled probe AUROC cannot establish within-attempt information without a question-only baseline or within-problem evaluation.

Abstract

Reasoning traces of large language models are widely read as containing "breakthrough" moments and early-legible fates. Both readings rest on measurements missing a counterfactual control at the level of the claim; we supply both controls. First, a restart-controlled truncation probe separates when a solution fits the continuation budget from when a prefix carries value that fresh computation cannot buy, comparing per-anchor continuation solve rates against from-scratch restart curves at matched total generated-token budget. Applied to 178 problem-model cells (89 MATH problems x two small open models, an outcome-blind but difficulty-targeted cohort), exactly 1 of 178 cells survives as prefix-limited; restart dose-response separates a compute-starved model from a capability-limited one; and wherever the matched budget lies inside the restart grid, continuing the model's own prefix beats restarting (9 of 9) -- predominantly compute compression rather than expanded reachability. Second, a pre-registered, difficulty-controlled test finds no detectable outcome information in early-window internal signals beyond a problem-difficulty baseline, and two generation-free analyses of public corpora show why this control is needed: a trace-blind difficulty proxy reaches AUROC 0.873 on 192K DeepSeek-R1 generations -- inside the published probe range -- and a closely matched reconstruction of the closest published early-window positive recovers a comparable pooled result (0.849) while within problem it is statistically indistinguishable from chance at all ten anchors (0.496 at t=4); a post-hoc within-targeted probe finds only a small average residual, concentrated in three low-failure problems. High pooled probe AUROCs cannot by themselves establish within-attempt information; a question-only baseline or within-problem evaluation is required.

Every Monday
Get next week’s papers.
Subscribe on Substack