🚀NEW LABGetting Started with Claude AgentsStart lab
Training

Automatic Gradient Descent: Deep Learning without Hyperparameters

First page
Automatic Gradient Descent: Deep Learning without Hyperparameters
Paper summary

A hyperparameter-free first-order optimizer that leverages architecture.

Ask this paper

Key points
01

Architecture-aware optimization: Derives optimization algorithms that explicitly account for neural network architecture rather than treating it as a black box.

02

No hyperparameters: Eliminates learning rate tuning - a hyperparameter-free optimizer that just works.

03

ImageNet scale: Successfully trains CNNs at ImageNet scale, demonstrating the approach scales to realistic workloads.

04

Optimizer research: Contributes to the ongoing search for optimizers that reduce tuning burden, complementing Adam-era hyperparameter-heavy methods.

Every Monday
Get next week’s papers.
Subscribe on Substack