AltUp (Alternating Updates)
First page

Paper summary
Google's AltUp lets transformers benefit from wider representations without paying the full compute cost at every layer.
Ask this paper
01
Wide-but-cheap representation: Widens the learned representation but only actively updates one sub-block per layer, leaving others untouched during that forward pass.
02
Predict-and-correct: A predict-and-correct mechanism updates the inactive sub-blocks with predictions, so they remain coherent without full computation.
03
Negligible latency increase: Achieves wider representations at negligible latency cost compared to matched-width dense transformers.
04
Scaling lever: Provides a middle-ground between narrow dense models and sparse MoE - wider without routing complexity.