🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety

Studying LLM Generalization with Influence Functions

Free while signed in. Answers cite the passages they came from.

First page
Studying LLM Generalization with Influence Functions
The curator’s take

Anthropic scales influence functions to LLMs up to 52B parameters to investigate generalization patterns.

Key points
01

Efficient scaling: Introduces computational tricks that make influence-function analysis tractable on LLMs with up to 52 billion parameters - a massive scale-up from prior work.

02

Cross-lingual generalization: Finds evidence of cross-lingual generalization, where training examples in one language influence predictions in another.

03

Middle-layer abstraction: Middle layers of the network appear responsible for the most abstract generalization patterns, supporting emerging interpretability narratives.

04

Alignment implications: Influence-function analysis gives alignment researchers a new tool for understanding which training data drives which model behaviors.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack