Chain-of-Thought Is Not Explainability
Free while signed in. Answers cite the passages they came from.

It challenges the common assumption that chain-of-thought (CoT) reasoning in LLMs is synonymous with interpretability. While CoT improves performance and offers a seemingly transparent rationale, the authors argue it is neither necessary nor sufficient for faithful explanation. Through a review of empirical evidence and mechanistic insights, the paper makes the case that CoT often diverges from the internal computations that actually drive model predictions.
Unfaithful reasoning is systematic: CoT rationales are often unfaithful to the underlying model computations. Examples include models silently correcting mistakes, being influenced by prompt biases, and using latent shortcuts while offering post-hoc rationales that omit these factors.
Widespread misuse in the literature: Of 1,000 CoT-focused papers surveyed, 24.4% explicitly use CoT as an interpretability technique without proper justification. The paper introduces a misclaim detection pipeline to track this trend and shows that the prevalence has not declined over time.
Architectural mismatch: Transformer models compute in a distributed, parallel manner that doesn’t align with the sequential nature of CoT explanations. This leads to plausible but misleading narratives that fail to capture causal dependencies.
Recommendations for improvement: The authors advocate for causal validation techniques (e.g., counterfactual interventions, activation patching), cognitive science-inspired mechanisms (e.g., metacognition, self-correction), and human-centered interfaces to assess CoT faithfulness. However, even these remain partial solutions to a deeper architectural challenge.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack