Bytes Are All You Need
First page

Paper summary
Performs classification directly on file bytes without decoding.
Ask this paper
01
Raw-byte input: Trains Transformers directly on raw file bytes (PNG, WAV, etc.) rather than decoded tensors.
02
Strong results: Achieves 77.33% ImageNet Top-1 accuracy on raw bytes and 95.42% on raw WAV for Speech Commands v2.
03
Format-agnostic: A single architecture handles any file format without preprocessing pipelines.
04
Infrastructure simplification: Suggests a future where models eat raw bytes and skip format-specific codecs - simpler pipelines with less preprocessing error.