No Positional Encodings (NoPE)
Free while signed in. Answers cite the passages they came from.

Shows explicit position embeddings aren't essential for decoder-only Transformers.
Implicit positional learning: Decoder-only Transformers learn positional information from the causal attention mask alone - no explicit encoding needed.
Length generalization: NoPE generalizes better to longer sequences than ALiBi and Rotary, which have surprising length-generalization issues.
Architectural simplification: Removing positional encodings simplifies the architecture with no quality loss on standard tasks.
Long-context influence: Informed the 2024 resurgence of interest in length-generalization-friendly architectures.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack