🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency

Medusa

First page
Medusa
Paper summary

Medusa accelerates LLM inference by bolting on multiple decoding heads that predict several future tokens in parallel, dramatically reducing decoding steps.

Ask this paper

Key points
01

Parallel multi-head decoding: Attaches K extra "Medusa heads" to a base LLM, each trained to predict the k-th next token; a verifier step accepts the longest consistent prefix.

02

2.2x+ speedup (Medusa-1): Delivers over 2.2x end-to-end inference speedup without quality loss when the heads are trained on top of a frozen backbone.

03

2.3-3.6x speedup (Medusa-2): Jointly fine-tuning the base model with the heads gives a further bump to 2.3-3.6x speedup while preserving generation quality.

04

Simpler than speculative decoding: Avoids running a separate draft model, making Medusa cheaper to deploy and easier to reason about than traditional speculative-decoding setups.

Every Monday
Get next week’s papers.
Subscribe on Substack