The Unreasonable Ineffectiveness of the Deeper Layers
Free while signed in. Answers cite the passages they came from.

The paper shows that open-weight LLMs tolerate removing up to half of their transformer blocks with only minor degradation, provided a short QLoRA pass is used to heal the damage afterwards.
Similarity-based pruning: A layer-similarity metric identifies contiguous blocks that can be removed without dramatically changing the hidden-state distribution at that depth.
Heal with QLoRA: After pruning, a small fine-tune with QLoRA on a single 40GB GPU recovers most of the lost performance - the whole recipe is reproducible on modest hardware.
Up to half the layers gone: Performance stays close to the original until a large fraction of layers (up to ~50%) has been removed, at which point accuracy drops sharply.
Implications: Either current pre-training underutilizes deeper layers or shallow layers carry most of the useful knowledge - either reading questions how efficiently today's LLMs use their depth.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack