🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency

Resource-efficient LLMs & Multimodal Foundation Models

First page
Resource-efficient LLMs & Multimodal Foundation Models
Paper summary

A wide-ranging survey of efficiency techniques for LLMs and multimodal foundation models, spanning architecture, algorithms, and system design.

Ask this paper

Key points
01

Three-pillar view: Organizes the space by architecture-level techniques (attention variants, MoE, SSMs), algorithm-level techniques (quantization, pruning, distillation), and system-level techniques (serving, scheduling, hardware).

02

Multimodal scope: Explicitly includes multimodal foundation models, not just text-only LLMs, covering vision-language and other modality combinations.

03

Practical designs: Connects research techniques to real deployment patterns - inference serving, batch scheduling, and memory-optimized training.

04

Benchmark reference: Aggregates numbers across techniques and models to give practitioners a single table for comparing efficiency trade-offs.

Every Monday
Get next week’s papers.
Subscribe on Substack