Wanda
First page

Paper summary
A simple, effective pruning approach for LLMs requiring no retraining.
Ask this paper
01
Weight×activation pruning: Prunes weights with the smallest magnitude × corresponding input activations on a per-output basis.
02
Zero retraining: Requires no retraining or weight updates, making it immediately deployable.
03
Simple beats complex: Outperforms magnitude-only pruning and matches or exceeds more complex training-based pruning methods.
04
Production pruning: Became a widely-adopted baseline in LLM pruning research due to its simplicity and strong performance.