Mixture-of-Depths

DeepMind proposes dynamically allocating transformer FLOPs across sequence positions via a top-k router, so "easy" tokens skip expensive blocks while "hard" tokens get full computation.
Ask this paper
Top-k routing: Each MoD layer picks a fixed-size subset of tokens to process, keeping the compute graph statically shaped while introducing conditional depth.
Matching baselines at lower FLOPs: MoD models match standard transformers on training loss for equivalent FLOPs, and outperform them at matched step count by spending compute where it matters.
Faster sampling: Inference gets up to 50% faster because many tokens bypass the heaviest layers without sacrificing downstream quality.
Static-compute advantage: Because the total per-batch FLOPs are predictable, MoD is easier to deploy on accelerators than token-routed MoE, combining the efficiency benefits of sparsity with dense execution patterns.