🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency · Architecture

The Efficiency Spectrum of LLMs

First page
The Efficiency Spectrum of LLMs
Paper summary

A comprehensive review of algorithmic advancements for improving LLM efficiency across the full training-to-inference stack.

Ask this paper

Key points
01

Scaling laws and data: Covers how scaling laws and data-utilization strategies interact with efficiency - more isn't always better under compute constraints.

02

Architectural innovations: Reviews attention variants, state-space models, MoE, and other architectural levers for efficient scaling.

03

Training and tuning: Catalogs PEFT methods (LoRA, adapters, prefix tuning), quantization-aware training, and curriculum-based training strategies.

04

Inference techniques: Surveys quantization, pruning, speculative decoding, KV-cache optimization, and batching as the inference-time efficiency toolkit.

Every Monday
Get next week’s papers.
Subscribe on Substack