Compression Algorithms for LLMs
Free while signed in. Answers cite the passages they came from.

A survey covering the main families of LLM compression techniques and when each one is appropriate.
Six technique families: Pruning, quantization, knowledge distillation, low-rank approximation, parameter sharing, and efficient architecture design - with worked examples in each.
Trade-off framing: Walks through memory, latency, and accuracy trade-offs so practitioners can pick techniques that match their deployment constraints.
Training vs. post-training: Distinguishes methods that require retraining or fine-tuning from purely post-training approaches, a practical axis for production teams.
Open problems: Flags hardware co-design, combining compression techniques, and evaluating compressed LLMs on long-context and reasoning tasks as under-explored areas.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack