MambaByte
Free while signed in. Answers cite the passages they came from.

MambaByte adapts the Mamba state-space architecture to learn directly from raw bytes, bypassing tokenization and all its well-known failure modes.
Token-free modeling: Works at the byte level, eliminating issues like tokenizer inefficiency, cross-lingual bias, and brittle handling of typos, whitespace, and code.
SSM advantage: Byte-level modeling normally blows up autoregressive transformer cost because sequences get much longer - Mamba's linear-time state-space scan sidesteps this directly.
Outperforms subword transformers: At matched compute, MambaByte outperforms subword-based transformer baselines on standard language-modeling benchmarks.
Fast inference: Reports substantial inference speedups over byte-level transformers, showing that SSMs may be the natural backbone for tokenization-free modeling.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack