🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture

No Positional Encodings (NoPE)

First page
No Positional Encodings (NoPE)
Paper summary

Shows explicit position embeddings aren't essential for decoder-only Transformers.

Ask this paper

Key points
01

Implicit positional learning: Decoder-only Transformers learn positional information from the causal attention mask alone - no explicit encoding needed.

02

Length generalization: NoPE generalizes better to longer sequences than ALiBi and Rotary, which have surprising length-generalization issues.

03

Architectural simplification: Removing positional encodings simplifies the architecture with no quality loss on standard tasks.

04

Long-context influence: Informed the 2024 resurgence of interest in length-generalization-friendly architectures.

Every Monday
Get next week’s papers.
Subscribe on Substack