🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Multimodal

Hierarchical Vision Transformer (Hiera)

First page
Hierarchical Vision Transformer (Hiera)
Paper summary

Pretrains ViTs with MAE while removing unnecessary multi-stage complexity.

Ask this paper

Key points
01

Simplified architecture: Strips away hand-designed components (shifted windows, relative position biases) from hierarchical ViTs like Swin.

02

MAE pretraining: Leverages masked autoencoder pretraining to compensate for reduced inductive bias.

03

Faster and more accurate: Achieves better accuracy and faster inference/training than prior hierarchical ViTs.

04

Architecture minimalism: Reinforces the "bitter lesson" direction - simpler architectures with better pretraining beat complex hand-designed ones.

Every Monday
Get next week’s papers.
Subscribe on Substack