The Larger They Are, the Harder They Fail
Free while signed in. Answers cite the passages they came from.

Reveals inverse-scaling failures in LLM code generation.
Function-name swap test: Swaps default Python function names and observes that larger LLMs fail harder to adapt - they prefer memorized patterns.
Inverse scaling: Counter to the usual "bigger is better" narrative, larger models prefer incorrect memorized continuations more strongly than smaller ones.
Memorization vs. reasoning: Highlights the tension between memorization (which helps on training data) and reasoning (which helps on novel data).
Safety implications: Important for safety/robustness - bigger models may be more brittle in adversarial or out-of-distribution settings.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack