Parallel Speculative Sampling
First page

Paper summary
Amazon researchers propose a parallel variant of speculative sampling that achieves significant LLM inference speedups with minimal extra parameters.
Ask this paper
01
Parallel decoding: Combines speculative sampling with parallel decoding so multiple tokens can be generated and verified in a single pass.
02
Tiny overhead: Requires learning only O(d_emb) additional parameters, far fewer than typical speculative-decoding draft models.
03
Up to 30% speedup: Achieves up to 30% end-to-end inference speedup without compromising output quality.
04
Minimal integration cost: Unlike separate-draft-model speculative decoding, this fits inside the main model with essentially no deployment overhead.