Skaling
Free while signed in. Answers cite the passages they came from.

Standard neural scaling laws assume model size and training data act on loss independently. That assumption bakes in a cross-derivative of exactly zero, and it is why the Chinchilla form drifts at the data-scarce and heavy-overtraining edges of the grid, which is exactly where deployment now happens.
One extra parameter, one coupling: The Skaling law generalizes the Chinchilla form by coupling capacity and data through a single interaction exponent, restoring the interaction that the additive form discards while adding only one parameter.
Errors shrink where they were worst: The extra term reduces mean absolute percentage error by 1.5x to 3x across both interpolation and extrapolation, and Skaling wins on 76% of configurations with a median improvement of 2.2x. The largest corrections land in the corners where standard laws show a saddle-shaped residual.
Cheaper profiling grids: Paired with an L-shape sparse grid restricted to low-compute runs, sweeping data volume for small models and model size at a fixed small data budget, it extrapolates the full grid using roughly 10x less compute than a uniform sweep.
Why it matters: Pretraining budgets are planned from fits to small runs, so a functional form that stays accurate past compute optimal and can be fit cheaply changes how those decisions get made. The empirical gradient analysis showing a real N-D interaction is the part worth reading closely.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack