LLM Self-Explanations
First page

Paper summary
Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.
Ask this paper
01
Self-explanation capability: LLMs can self-generate feature-attribution explanations that meaningfully highlight the tokens driving their predictions.
02
Performance + truthfulness: Self-explanation improves both task performance and the truthfulness of outputs compared to baseline prompting.
03
CoT synergy: Combines productively with chain-of-thought prompting, giving additive improvements rather than substituting for it.
04
Interpretability lever: Offers a cheap, model-agnostic interpretability pattern that works through the API without needing gradients or white-box access.