MEGABYTE
First page

Paper summary
Multiscale Transformers for predicting million-byte sequences.
Ask this paper
01
Two-level architecture: Combines a large global Transformer over patches with a smaller local Transformer over bytes within each patch.
02
Sub-quadratic attention: Achieves sub-quadratic self-attention cost through the patch-level hierarchy, enabling million-byte sequences.
03
Decoding parallelism: Improves decoding parallelism compared to flat Transformers that must decode token-by-token.
04
Tokenization-free: Operates directly on bytes without tokenizers - potentially avoiding tokenizer failure modes.