LLM Self-Explanations
Free while signed in. Answers cite the passages they came from.

Investigates whether LLMs can generate useful feature-attribution explanations for their own outputs.
Self-explanation capability: LLMs can self-generate feature-attribution explanations that meaningfully highlight the tokens driving their predictions.
Performance + truthfulness: Self-explanation improves both task performance and the truthfulness of outputs compared to baseline prompting.
CoT synergy: Combines productively with chain-of-thought prompting, giving additive improvements rather than substituting for it.
Interpretability lever: Offers a cheap, model-agnostic interpretability pattern that works through the API without needing gradients or white-box access.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack