Scaling Transformer to 1M tokens with RMT
First page

Paper summary
Recurrent Memory Transformer extends BERT's effective context to 2M tokens.
Ask this paper
01
Recurrent memory mechanism: Augments BERT with a recurrent memory that carries information across segments, enabling massive context lengths.
02
2M token context: Scales effective context to two million tokens while maintaining high memory retrieval accuracy.
03
Segment-level recurrence: Processes input in segments while passing a compressed memory token stream across them.
04
Long-context trend: Part of the 2023 explosion of long-context techniques that established ultra-long context as a viable research direction.