🚀NEW LABGetting Started with Claude AgentsStart lab
Training

Chinchilla Scaling: A replication attempt

First page
Chinchilla Scaling: A replication attempt
Paper summary

This paper re-examines the third estimation procedure in Hoffmann et al. (2022) Chinchilla scaling law and finds it is inconsistent with the paper's own first two methods, fails to fit the extracted data, and reports implausibly narrow confidence intervals.

Ask this paper

Key points
01

What was audited: Chinchilla proposed three independent methods to estimate the compute-optimal ratio of parameters to training tokens; this paper digs into the parametric-loss-fitting approach (method 3).

02

Inconsistent estimates: The published estimates from method 3 do not match the predictions of methods 1 and 2, and the parametric fit does not actually pass through the reconstructed data points.

03

Implausible confidence intervals: The reported intervals would statistically require over 600,000 training runs, whereas the authors likely ran fewer than 500 - suggesting methodological errors in the uncertainty quantification.

04

Rederivation: A corrected fit using method 3 produces scaling estimates that are consistent with methods 1 and 2, restoring internal coherence and slightly revising the compute-optimal guidance.

Every Monday
Get next week’s papers.
Subscribe on Substack