🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Efficiency

MEGABYTE

First page
MEGABYTE
Paper summary

Multiscale Transformers for predicting million-byte sequences.

Ask this paper

Key points
01

Two-level architecture: Combines a large global Transformer over patches with a smaller local Transformer over bytes within each patch.

02

Sub-quadratic attention: Achieves sub-quadratic self-attention cost through the patch-level hierarchy, enabling million-byte sequences.

03

Decoding parallelism: Improves decoding parallelism compared to flat Transformers that must decode token-by-token.

04

Tokenization-free: Operates directly on bytes without tokenizers - potentially avoiding tokenizer failure modes.

Every Monday
Get next week’s papers.
Subscribe on Substack