🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture

AltUp (Alternating Updates)

First page
AltUp (Alternating Updates)
Paper summary

Google's AltUp lets transformers benefit from wider representations without paying the full compute cost at every layer.

Ask this paper

Key points
01

Wide-but-cheap representation: Widens the learned representation but only actively updates one sub-block per layer, leaving others untouched during that forward pass.

02

Predict-and-correct: A predict-and-correct mechanism updates the inactive sub-blocks with predictions, so they remain coherent without full computation.

03

Negligible latency increase: Achieves wider representations at negligible latency cost compared to matched-width dense transformers.

04

Scaling lever: Provides a middle-ground between narrow dense models and sparse MoE - wider without routing complexity.

Every Monday
Get next week’s papers.
Subscribe on Substack