🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training

Skaling

Free while signed in. Answers cite the passages they came from.

First page
Skaling
The curator’s take

Standard neural scaling laws assume model size and training data act on loss independently. That assumption bakes in a cross-derivative of exactly zero, and it is why the Chinchilla form drifts at the data-scarce and heavy-overtraining edges of the grid, which is exactly where deployment now happens.

Key points
01

One extra parameter, one coupling: The Skaling law generalizes the Chinchilla form by coupling capacity and data through a single interaction exponent, restoring the interaction that the additive form discards while adding only one parameter.

02

Errors shrink where they were worst: The extra term reduces mean absolute percentage error by 1.5x to 3x across both interpolation and extrapolation, and Skaling wins on 76% of configurations with a median improvement of 2.2x. The largest corrections land in the corners where standard laws show a saddle-shaped residual.

03

Cheaper profiling grids: Paired with an L-shape sparse grid restricted to low-compute runs, sweeping data volume for small models and model size at a fixed small data budget, it extrapolates the full grid using roughly 10x less compute than a uniform sweep.

04

Why it matters: Pretraining budgets are planned from fits to small runs, so a functional form that stays accurate past compute optimal and can be fit cheaply changes how those decisions get made. The empirical gradient analysis showing a real N-D interaction is the part worth reading closely.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack