🚀NEW LABGetting Started with Claude AgentsStart lab
Architecture · Memory

Frontier Models Struggle to Copy

First page
Frontier Models Struggle to Copy
Paper summary

Frontier models can write proofs yet stumble on faithfully copying a long block of text that sits well within their context window. This paper traces the failure to 1D positional encodings, whose inductive bias favors a copying shortcut based on matching local context rather than carefully locating the corresponding input positions. The fix is 2D-RoPE, which lays text out on a 2D grid and gives each token a row and a column ID, so copying becomes retrieving tokens at a fixed column offset. Shallow Transformers with 2D-RoPE copy perfectly at input lengths hundreds of times longer than those seen in training.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack