AltUp (Alternating Updates)
Free while signed in. Answers cite the passages they came from.

Google's AltUp lets transformers benefit from wider representations without paying the full compute cost at every layer.
Wide-but-cheap representation: Widens the learned representation but only actively updates one sub-block per layer, leaving others untouched during that forward pass.
Predict-and-correct: A predict-and-correct mechanism updates the inactive sub-blocks with predictions, so they remain coherent without full computation.
Negligible latency increase: Achieves wider representations at negligible latency cost compared to matched-width dense transformers.
Scaling lever: Provides a middle-ground between narrow dense models and sparse MoE - wider without routing complexity.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack