LOMO
First page

Paper summary
A memory-efficient optimizer that combines gradient computation and parameter update in one step.
Ask this paper
01
Fused grad-update: Fuses backpropagation and SGD update into a single operation, eliminating the need to store all gradients in memory simultaneously.
02
Full-parameter tuning: Enables full-parameter fine-tuning of a 65B LLM on a single 8x24GB GPU machine.
03
Democratization: Makes full fine-tuning (not just LoRA) accessible to researchers without multi-node GPU clusters.
04
Optimizer memory research: Joined the 2023 wave of optimizer memory innovations (8-bit Adam, AdaFactor, GaLore) democratizing large-model tuning.