Simplifying Transformer Blocks
Free while signed in. Answers cite the passages they came from.

Researchers show that many components of the standard transformer block can be removed with no loss in training speed or quality.
Aggressive simplification: Removes residual connections, normalization layers, and value/projection parameters in specific blocks without hurting per-update training speed.
Works across architectures: Tested on autoregressive decoder-only and BERT encoder-only models, validating that the simplifications aren't architecture-specific.
15% faster throughput: Simplified blocks deliver 15% faster training throughput with fewer parameters - a clean efficiency win.
Design-space implication: Suggests the standard transformer is overdetermined and that careful ablation can yield simpler, faster architectures without new ideas.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack