The Efficiency Spectrum of LLMs
Free while signed in. Answers cite the passages they came from.

A comprehensive review of algorithmic advancements for improving LLM efficiency across the full training-to-inference stack.
Scaling laws and data: Covers how scaling laws and data-utilization strategies interact with efficiency - more isn't always better under compute constraints.
Architectural innovations: Reviews attention variants, state-space models, MoE, and other architectural levers for efficient scaling.
Training and tuning: Catalogs PEFT methods (LoRA, adapters, prefix tuning), quantization-aware training, and curriculum-based training strategies.
Inference techniques: Surveys quantization, pruning, speculative decoding, KV-cache optimization, and batching as the inference-time efficiency toolkit.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack