🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Efficiency · Training

LOMO

Free while signed in. Answers cite the passages they came from.

First page
LOMO
The curator’s take

A memory-efficient optimizer that combines gradient computation and parameter update in one step.

Key points
01

Fused grad-update: Fuses backpropagation and SGD update into a single operation, eliminating the need to store all gradients in memory simultaneously.

02

Full-parameter tuning: Enables full-parameter fine-tuning of a 65B LLM on a single 8x24GB GPU machine.

03

Democratization: Makes full fine-tuning (not just LoRA) accessible to researchers without multi-node GPU clusters.

04

Optimizer memory research: Joined the 2023 wave of optimizer memory innovations (8-bit Adam, AdaFactor, GaLore) democratizing large-model tuning.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack