The Boundary of Neural Network Trainability is Fractal

Sohl-Dickstein finds that the boundary between trainable and untrainable hyperparameter configurations looks like a Mandelbrot-style fractal across many architectures.
Ask this paper
Fractal boundaries: Zooming into the hyperparameter plane (e.g. learning rate × init scale), the trainable region has a self-similar boundary that stays fractal over more than ten decades of zoom.
Universal across setups: Observed for every tested neural network configuration including deep linear networks, suggesting the fractal structure is a property of the training dynamics itself.
Edge-of-stability: The best-performing hyperparameter choices tend to sit right at the edge of stability, consistent with recent "edge-of-stability" observations in deep learning theory.
Implication: Hyperparameter search is more analogous to exploring a chaotic dynamical system than a smooth optimization surface, which has implications for both practitioners and learning theorists.