Transformers without Normalization

Researchers from Meta, NYU, MIT, and Princeton present a surprisingly simple method, Dynamic Tanh (DyT), that removes normalization layers (e.g. LayerNorm, RMSNorm) in Transformers while achieving equal or better results. Key ideas include:
Ask this paper
Tanh-like mapping of LayerNorm β By analyzing trained models, they observe that LayerNorm often behaves like an S-shaped tanh function, scaling inputs while squashing extremes.
Dynamic Tanh (DyT) β Replaces each normalization layer with a per-channel tanh(Ξ±x) and learnable affine parameters. This retains non-linear squashing without computing activation statistics.
Stable convergence, on par with LN β Across tasks (vision, speech, diffusion, language modeling), DyT-based models match or exceed normalized baselines without extra tuning. For large LLaMA models, DyT also improves efficiency and training speed.
Efficient, widely applicable β Eliminating normalization operations saves computation overhead. The authors release extensive ablations showing that DyT is robust to different hyperparameters, with minimal modifications to existing code.