🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning · Safety

Unfaithful Explanations in Chain-of-Thought Prompting

Free while signed in. Answers cite the passages they came from.

First page
Unfaithful Explanations in Chain-of-Thought Prompting
The curator’s take

Demonstrates CoT explanations can misrepresent the true reason for a model's prediction.

Key points
01

Biased-CoT demonstration: Shows when models are biased toward incorrect answers (e.g., from few-shot bias), they generate CoT justifications supporting those wrong answers.

02

Confident-but-wrong: The CoT sounds plausible and confident even when it's post-hoc rationalization rather than actual reasoning.

03

Interpretability warning: An important caution that visible reasoning traces shouldn't be uncritically trusted as explanations.

04

Safety implications: Part of the growing evidence base that CoT monitoring for safety has limitations.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack