🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Studying LLM Generalization with Influence Functions

First page
Studying LLM Generalization with Influence Functions
Paper summary

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

Ask this paper

Key points
01

Efficient scaling: Introduces computational tricks that make influence-function analysis tractable on LLMs with up to 52 billion parameters - a massive scale-up from prior work.

02

Cross-lingual generalization: Finds evidence of cross-lingual generalization, where training examples in one language influence predictions in another.

03

Middle-layer abstraction: Middle layers of the network appear responsible for the most abstract generalization patterns, supporting emerging interpretability narratives.

04

Alignment implications: Influence-function analysis gives alignment researchers a new tool for understanding which training data drives which model behaviors.

Every Monday
Get next week’s papers.
Subscribe on Substack