🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Memory

Scaling Transformer to 1 Billion Tokens (LongNet)

Free while signed in. Answers cite the passages they came from.

First page
Scaling Transformer to 1 Billion Tokens (LongNet)
The curator’s take

Microsoft's Transformer variant scaling sequence length past 1B tokens.

Key points
01

Dilated attention: Introduces dilated attention that exponentially grows the attention field, enabling linear complexity in sequence length.

02

No short-sequence loss: Achieves extreme long-context scaling with no degradation on shorter sequences.

03

1B token demo: Demonstrates viability at the 1-billion token context scale - an order of magnitude beyond anything previously attempted.

04

Long-context frontier: Pushed the frontier of what's theoretically possible for ultra-long-context Transformers, even though production models stayed in the hundreds-of-thousands-of-tokens range.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack