πŸš€NEW LABGetting Started with Claude AgentsStart lab
Architecture

Transformers without Normalization

First page
Transformers without Normalization
Paper summary

Researchers from Meta, NYU, MIT, and Princeton present a surprisingly simple method, Dynamic Tanh (DyT), that removes normalization layers (e.g. LayerNorm, RMSNorm) in Transformers while achieving equal or better results. Key ideas include:

Ask this paper

Key points
01

Tanh-like mapping of LayerNorm – By analyzing trained models, they observe that LayerNorm often behaves like an S-shaped tanh function, scaling inputs while squashing extremes.

02

Dynamic Tanh (DyT) – Replaces each normalization layer with a per-channel tanh(Ξ±x) and learnable affine parameters. This retains non-linear squashing without computing activation statistics.

03

Stable convergence, on par with LN – Across tasks (vision, speech, diffusion, language modeling), DyT-based models match or exceed normalized baselines without extra tuning. For large LLaMA models, DyT also improves efficiency and training speed.

04

Efficient, widely applicable – Eliminating normalization operations saves computation overhead. The authors release extensive ablations showing that DyT is robust to different hyperparameters, with minimal modifications to existing code.

Every Monday
Get next week’s papers.
Subscribe on Substack