Measuring Faithfulness in Chain-of-Thought Reasoning
Paper preview

Paper summary
Anthropic's investigation into whether CoT reasoning actually reflects the model's internal decision process.
Ask this paper
01
Intervention protocol: Uses paraphrasing, mistake-injection, and truncation of reasoning chains to test whether final answers depend on the visible reasoning.
02
Inverse scaling finding: Demonstrates that as models get larger and more capable, the reasoning becomes less faithful - an important inverse-scaling signal.
03
Task variability: Faithfulness varies significantly across tasks; some tasks/model-sizes support CoT that is meaningfully tied to the answer.
04
Interpretability foundation: Influential for subsequent interpretability and safety work on whether chain-of-thought can be trusted for monitoring model reasoning.