Unfaithful Explanations in Chain-of-Thought Prompting
First page

Paper summary
Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.
Ask this paper
01
Biased-CoT demonstration: Shows when models are biased toward incorrect answers (e.g., from few-shot bias), they generate CoT justifications supporting those wrong answers.
02
Confident-but-wrong: The CoT sounds plausible and confident even when it's post-hoc rationalization rather than actual reasoning.
03
Interpretability warning: An important caution that visible reasoning traces shouldn't be uncritically trusted as explanations.
04
Safety implications: Part of the growing evidence base that CoT monitoring for safety has limitations.