Unfaithful Explanations in Chain-of-Thought Prompting
Free while signed in. Answers cite the passages they came from.

Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.
Biased-CoT demonstration: Shows when models are biased toward incorrect answers (e.g., from few-shot bias), they generate CoT justifications supporting those wrong answers.
Confident-but-wrong: The CoT sounds plausible and confident even when it's post-hoc rationalization rather than actual reasoning.
Interpretability warning: An important caution that visible reasoning traces shouldn't be uncritically trusted as explanations.
Safety implications: Part of the growing evidence base that CoT monitoring for safety has limitations.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack