Model Compression for LLMs Survey
Free while signed in. Answers cite the passages they came from.

A survey of recent model-compression techniques applied specifically to LLMs.
Core technique families: Covers quantization, pruning, knowledge distillation, and architectural compression across training-time and post-training approaches.
LLM-specific concerns: Addresses unique LLM concerns including long-sequence compression, KV-cache optimization, and retaining reasoning capability under compression.
Evaluation metrics: Reviews benchmark strategies and evaluation metrics for measuring compressed-LLM effectiveness - not just perplexity but downstream capability preservation.
Practitioner reference: Functions as a compact reference for teams deciding which compression technique matches their deployment constraints.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack