Ring Attention
First page

Paper summary
UC Berkeley's Ring Attention scales transformer context to 100M+ tokens by distributing blockwise self-attention across devices in a ring topology.
Ask this paper
01
Blockwise attention: Computes self-attention in blocks so that only small KV chunks need to fit on each device at any time.
02
Ring communication: Passes KV chunks between devices in a ring, overlapping communication with computation to hide networking latency.
03
Context scales with devices: Achievable context length grows linearly with the number of devices, with no attention approximations required.
04
100M+ tokens: Enables context lengths exceeding 100 million tokens in theory, far beyond what any single-device attention implementation can reach.