🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 4 – Sep 4, 2026
Reasoning · Safety

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

First page
Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
The curator’s take

Kevin Du, Alexander Hoyle, Laura Ruis and Acyr Locatelli (ETH Zurich, work done during an internship at Cohere) test whether the text of a reasoning step actually encodes how much that step mattered, using Monte Carlo advantage as ground truth, and find only partial recoverability.

Ask this paper

Key points
01

A concrete definition of importance: a step's advantage, the change in expected reward from including it, estimated via Monte Carlo rollouts, which replaces judge-vs-judge circularity with a measurable target.

02

Judges beat prevalence but not by much: sufficiently capable LLM judges outperform a prevalence baseline yet fall well short of the noise ceiling on identifying high-advantage steps.

03

Fine-tuning helps asymmetrically: a trained step-level critic improves strongly on incorrect responses but stays far from ceiling on correct ones, which is the harder and more common case in PRM training data.

04

The title is the thesis: legibility of a trace is not interpretability of it, and step importance is only partially recoverable from step text.

05

Direct implication for process reward modeling: PRMs and generative critics are supervised on a signal the text only partly carries, which bounds how good they can get without changing the input.

Abstract

Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.

Every Monday
Get next week’s papers.
Subscribe on Substack