🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Safety

Unfaithful Explanations in Chain-of-Thought Prompting

First page
Unfaithful Explanations in Chain-of-Thought Prompting
Paper summary

Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.

Ask this paper

Key points
01

Biased-CoT demonstration: Shows when models are biased toward incorrect answers (e.g., from few-shot bias), they generate CoT justifications supporting those wrong answers.

02

Confident-but-wrong: The CoT sounds plausible and confident even when it's post-hoc rationalization rather than actual reasoning.

03

Interpretability warning: An important caution that visible reasoning traces shouldn't be uncritically trusted as explanations.

04

Safety implications: Part of the growing evidence base that CoT monitoring for safety has limitations.

Every Monday
Get next week’s papers.
Subscribe on Substack