🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

LLM Self-Explanations

First page
LLM Self-Explanations
Paper summary

Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.

Ask this paper

Key points
01

Self-explanation capability: LLMs can self-generate feature-attribution explanations that meaningfully highlight the tokens driving their predictions.

02

Performance + truthfulness: Self-explanation improves both task performance and the truthfulness of outputs compared to baseline prompting.

03

CoT synergy: Combines productively with chain-of-thought prompting, giving additive improvements rather than substituting for it.

04

Interpretability lever: Offers a cheap, model-agnostic interpretability pattern that works through the API without needing gradients or white-box access.

Every Monday
Get next week’s papers.
Subscribe on Substack