MEGABYTE
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsMultiscale Transformers for predicting million-byte sequences.
01
Two-level architecture: Combines a large global Transformer over patches with a smaller local Transformer over bytes within each patch.
02
Sub-quadratic attention: Achieves sub-quadratic self-attention cost through the patch-level hierarchy, enabling million-byte sequences.
03
Decoding parallelism: Improves decoding parallelism compared to flat Transformers that must decode token-by-token.
04
Tokenization-free: Operates directly on bytes without tokenizers - potentially avoiding tokenizer failure modes.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack