Automatic Gradient Descent: Deep Learning without Hyperparameters
First page

Paper summary
A hyperparameter-free first-order optimizer that leverages architecture.
Ask this paper
01
Architecture-aware optimization: Derives optimization algorithms that explicitly account for neural network architecture rather than treating it as a black box.
02
No hyperparameters: Eliminates learning rate tuning - a hyperparameter-free optimizer that just works.
03
ImageNet scale: Successfully trains CNNs at ImageNet scale, demonstrating the approach scales to realistic workloads.
04
Optimizer research: Contributes to the ongoing search for optimizers that reduce tuning burden, complementing Adam-era hyperparameter-heavy methods.