Chinchilla Scaling: A replication attempt

This paper re-examines the third estimation procedure in Hoffmann et al. (2022) Chinchilla scaling law and finds it is inconsistent with the paper's own first two methods, fails to fit the extracted data, and reports implausibly narrow confidence intervals.
Ask this paper
What was audited: Chinchilla proposed three independent methods to estimate the compute-optimal ratio of parameters to training tokens; this paper digs into the parametric-loss-fitting approach (method 3).
Inconsistent estimates: The published estimates from method 3 do not match the predictions of methods 1 and 2, and the parametric fit does not actually pass through the reconstructed data points.
Implausible confidence intervals: The reported intervals would statistically require over 600,000 training runs, whereas the authors likely ran fewer than 500 - suggesting methodological errors in the uncertainty quantification.
Rederivation: A corrected fit using method 3 produces scaling estimates that are consistent with methods 1 and 2, restoring internal coherence and slightly revising the compute-optimal guidance.