The Unreasonable Ineffectiveness of the Deeper Layers

The paper shows that open-weight LLMs tolerate removing up to half of their transformer blocks with only minor degradation, provided a short QLoRA pass is used to heal the damage afterwards.
Ask this paper
Similarity-based pruning: A layer-similarity metric identifies contiguous blocks that can be removed without dramatically changing the hidden-state distribution at that depth.
Heal with QLoRA: After pruning, a small fine-tune with QLoRA on a single 40GB GPU recovers most of the lost performance - the whole recipe is reproducible on modest hardware.
Up to half the layers gone: Performance stays close to the original until a large fraction of layers (up to ~50%) has been removed, at which point accuracy drops sharply.
Implications: Either current pre-training underutilizes deeper layers or shallow layers carry most of the useful knowledge - either reading questions how efficiently today's LLMs use their depth.