🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Safety · Training

Explaining Grokking

Free while signed in. Answers cite the passages they came from.

First page
Explaining Grokking
The curator’s take

DeepMind advances our understanding of grokking, predicting and confirming two novel phenomena that test their theory.

Key points
01

Ungrokking: A model can go from perfect generalization back to memorization when trained further on a smaller dataset below a critical threshold - the first demonstration of this reverse effect.

02

Semi-grokking: A randomly initialized network trained on the critical dataset size shows a grokking-like transition but partial, rather than the sharp full-grokking curve.

03

Theoretical predictions: These behaviors were predicted from theory before being demonstrated empirically - a rare example of predictive rather than post-hoc explanation in deep learning.

04

Generalization theory: Advances understanding of when and why neural networks transition from memorization to generalization, bridging empirical observation with principled prediction.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack