Scaling Transformer to 1M tokens with RMT
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsRecurrent Memory Transformer extends BERT's effective context to 2M tokens.
01
Recurrent memory mechanism: Augments BERT with a recurrent memory that carries information across segments, enabling massive context lengths.
02
2M token context: Scales effective context to two million tokens while maintaining high memory retrieval accuracy.
03
Segment-level recurrence: Processes input in segments while passing a compressed memory token stream across them.
04
Long-context trend: Part of the 2023 explosion of long-context techniques that established ultra-long context as a viable research direction.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack