🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

The Hydra Effect

First page
The Hydra Effect
Paper summary

DeepMind shows that language models exhibit self-repairing behavior when attention heads are ablated.

Ask this paper

Key points
01

Self-repair phenomenon: Ablating a layer of attention heads causes a later layer to take over the ablated layer's function - a previously unknown redundancy property.

02

Interpretability implications: Complicates interpretability work based on ablation - removing a component doesn't necessarily isolate its contribution if other components compensate.

03

Circuit-level redundancy: Suggests transformer circuits have built-in redundancy that is activated under ablation, analogous to biological neural networks.

04

Research-method correction: Forces a rethinking of causal-mediation experiments in mechanistic interpretability, since ablations alone understate components' true contributions.

Every Monday
Get next week’s papers.
Subscribe on Substack