Resource-efficient LLMs & Multimodal Foundation Models
Free while signed in. Answers cite the passages they came from.

A wide-ranging survey of efficiency techniques for LLMs and multimodal foundation models, spanning architecture, algorithms, and system design.
Three-pillar view: Organizes the space by architecture-level techniques (attention variants, MoE, SSMs), algorithm-level techniques (quantization, pruning, distillation), and system-level techniques (serving, scheduling, hardware).
Multimodal scope: Explicitly includes multimodal foundation models, not just text-only LLMs, covering vision-language and other modality combinations.
Practical designs: Connects research techniques to real deployment patterns - inference serving, batch scheduling, and memory-optimized training.
Benchmark reference: Aggregates numbers across techniques and models to give practitioners a single table for comparing efficiency trade-offs.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack