🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Efficiency

Star Attention: Efficient LLM Inference over Long Sequences

First page
Star Attention: Efficient LLM Inference over Long Sequences
Paper summary

introduces Star Attention, a two-phase attention mechanism that processes long sequences by combining blockwise-local attention for context encoding with sequence-global attention for query processing and token generation; achieves up to 11x faster inference speeds while maintaining 95-100% accuracy compared to traditional attention mechanisms by efficiently distributing computation across multiple hosts; a key innovation is the "anchor block" mechanism, where each context block is prefixed with the first block, enabling effective approximation of global attention patterns while reducing computational overhead.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack