Scaling Transformer to 1 Billion Tokens (LongNet)
Free while signed in. Answers cite the passages they came from.

Microsoft's Transformer variant scaling sequence length past 1B tokens.
Dilated attention: Introduces dilated attention that exponentially grows the attention field, enabling linear complexity in sequence length.
No short-sequence loss: Achieves extreme long-context scaling with no degradation on shorter sequences.
1B token demo: Demonstrates viability at the 1-billion token context scale - an order of magnitude beyond anything previously attempted.
Long-context frontier: Pushed the frontier of what's theoretically possible for ultra-long-context Transformers, even though production models stayed in the hundreds-of-thousands-of-tokens range.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack