🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

Chinchilla Scaling: A replication attempt

Free while signed in. Answers cite the passages they came from.

First page
Chinchilla Scaling: A replication attempt
The curator’s take

This paper re-examines the third estimation procedure in Hoffmann et al. (2022) Chinchilla scaling law and finds it is inconsistent with the paper's own first two methods, fails to fit the extracted data, and reports implausibly narrow confidence intervals.

Key points
01

What was audited: Chinchilla proposed three independent methods to estimate the compute-optimal ratio of parameters to training tokens; this paper digs into the parametric-loss-fitting approach (method 3).

02

Inconsistent estimates: The published estimates from method 3 do not match the predictions of methods 1 and 2, and the parametric fit does not actually pass through the reconstructed data points.

03

Implausible confidence intervals: The reported intervals would statistically require over 600,000 training runs, whereas the authors likely ran fewer than 500 - suggesting methodological errors in the uncertainty quantification.

04

Rederivation: A corrected fit using method 3 produces scaling estimates that are consistent with methods 1 and 2, restoring internal coherence and slightly revising the compute-optimal guidance.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack