Studying LLM Generalization with Influence Functions

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.
Ask this paper
Efficient scaling: Introduces computational tricks that make influence-function analysis tractable on LLMs with up to 52 billion parameters - a massive scale-up from prior work.
Cross-lingual generalization: Finds evidence of cross-lingual generalization, where training examples in one language influence predictions in another.
Middle-layer abstraction: Middle layers of the network appear responsible for the most abstract generalization patterns, supporting emerging interpretability narratives.
Alignment implications: Influence-function analysis gives alignment researchers a new tool for understanding which training data drives which model behaviors.