🚀NEW LABGetting Started with Claude AgentsStart lab
Training

The Boundary of Neural Network Trainability is Fractal

First page
The Boundary of Neural Network Trainability is Fractal
Paper summary

Sohl-Dickstein finds that the boundary between trainable and untrainable hyperparameter configurations looks like a Mandelbrot-style fractal across many architectures.

Ask this paper

Key points
01

Fractal boundaries: Zooming into the hyperparameter plane (e.g. learning rate × init scale), the trainable region has a self-similar boundary that stays fractal over more than ten decades of zoom.

02

Universal across setups: Observed for every tested neural network configuration including deep linear networks, suggesting the fractal structure is a property of the training dynamics itself.

03

Edge-of-stability: The best-performing hyperparameter choices tend to sit right at the edge of stability, consistent with recent "edge-of-stability" observations in deep learning theory.

04

Implication: Hyperparameter search is more analogous to exploring a chaotic dynamical system than a smooth optimization surface, which has implications for both practitioners and learning theorists.

Every Monday
Get next week’s papers.
Subscribe on Substack