🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Efficiency · Architecture

The Efficiency Spectrum of LLMs

Free while signed in. Answers cite the passages they came from.

First page
The Efficiency Spectrum of LLMs
The curator’s take

A comprehensive review of algorithmic advancements for improving LLM efficiency across the full training-to-inference stack.

Key points
01

Scaling laws and data: Covers how scaling laws and data-utilization strategies interact with efficiency - more isn't always better under compute constraints.

02

Architectural innovations: Reviews attention variants, state-space models, MoE, and other architectural levers for efficient scaling.

03

Training and tuning: Catalogs PEFT methods (LoRA, adapters, prefix tuning), quantization-aware training, and curriculum-based training strategies.

04

Inference techniques: Surveys quantization, pruning, speculative decoding, KV-cache optimization, and batching as the inference-time efficiency toolkit.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack