Vision Transformers Need Registers
Free while signed in. Answers cite the passages they came from.

Meta researchers identify artifact tokens in ViT feature maps and propose a trivial fix: add dedicated register tokens.
Artifact identification: Vision transformers repurpose certain input tokens as "internal scratch space", producing high-norm artifacts that contaminate feature maps.
Register tokens: Adds a small number of dedicated register tokens to the input sequence, giving the model explicit scratch space instead of co-opting patch tokens.
Cleaner features: The fix produces substantially smoother feature and attention maps, with the artifact tokens disappearing.
New SoTA on dense tasks: Sets new state-of-the-art results on dense visual prediction tasks (segmentation, depth, object discovery), with real downstream impact.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack