Automatic Gradient Descent: Deep Learning without Hyperparameters
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsA hyperparameter-free first-order optimizer that leverages architecture.
01
Architecture-aware optimization: Derives optimization algorithms that explicitly account for neural network architecture rather than treating it as a black box.
02
No hyperparameters: Eliminates learning rate tuning - a hyperparameter-free optimizer that just works.
03
ImageNet scale: Successfully trains CNNs at ImageNet scale, demonstrating the approach scales to realistic workloads.
04
Optimizer research: Contributes to the ongoing search for optimizers that reduce tuning burden, complementing Adam-era hyperparameter-heavy methods.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack