🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency · Memory

MagicDec

First page
MagicDec
Paper summary

shows how speculative decoding can enhance throughput, reduce latency, and maintain accuracy in long context generation scenarios; it finds that as sequence length and batch size increase, bottlenecks shift from compute-bound to memory-bound; using these insights, they show it's possible to more effectively use speculative decoding for longer sequences, even when using large batch sizes.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack