Studying LLM Generalization with Influence Functions
Free while signed in. Answers cite the passages they came from.

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.
Efficient scaling: Introduces computational tricks that make influence-function analysis tractable on LLMs with up to 52 billion parameters - a massive scale-up from prior work.
Cross-lingual generalization: Finds evidence of cross-lingual generalization, where training examples in one language influence predictions in another.
Middle-layer abstraction: Middle layers of the network appear responsible for the most abstract generalization patterns, supporting emerging interpretability narratives.
Alignment implications: Influence-function analysis gives alignment researchers a new tool for understanding which training data drives which model behaviors.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack