The Larger They Are, the Harder They Fail
First page

Paper summary
Reveals inverse-scaling failures in LLM code generation.
Ask this paper
01
Function-name swap test: Swaps default Python function names and observes that larger LLMs fail harder to adapt - they prefer memorized patterns.
02
Inverse scaling: Counter to the usual "bigger is better" narrative, larger models prefer incorrect memorized continuations more strongly than smaller ones.
03
Memorization vs. reasoning: Highlights the tension between memorization (which helps on training data) and reasoning (which helps on novel data).
04
Safety implications: Important for safety/robustness - bigger models may be more brittle in adversarial or out-of-distribution settings.