Compression Algorithms for LLMs

A survey covering the main families of LLM compression techniques and when each one is appropriate.
Ask this paper
Six technique families: Pruning, quantization, knowledge distillation, low-rank approximation, parameter sharing, and efficient architecture design - with worked examples in each.
Trade-off framing: Walks through memory, latency, and accuracy trade-offs so practitioners can pick techniques that match their deployment constraints.
Training vs. post-training: Distinguishes methods that require retraining or fine-tuning from purely post-training approaches, a practical axis for production teams.
Open problems: Flags hardware co-design, combining compression techniques, and evaluating compressed LLMs on long-context and reasoning tasks as under-explored areas.