Simplifying Transformer Blocks

Researchers show that many components of the standard transformer block can be removed with no loss in training speed or quality.
Ask this paper
Aggressive simplification: Removes residual connections, normalization layers, and value/projection parameters in specific blocks without hurting per-update training speed.
Works across architectures: Tested on autoregressive decoder-only and BERT encoder-only models, validating that the simplifications aren't architecture-specific.
15% faster throughput: Simplified blocks deliver 15% faster training throughput with fewer parameters - a clean efficiency win.
Design-space implication: Suggests the standard transformer is overdetermined and that careful ablation can yield simpler, faster architectures without new ideas.