LOMO
Free while signed in. Answers cite the passages they came from.

A memory-efficient optimizer that combines gradient computation and parameter update in one step.
Fused grad-update: Fuses backpropagation and SGD update into a single operation, eliminating the need to store all gradients in memory simultaneously.
Full-parameter tuning: Enables full-parameter fine-tuning of a 65B LLM on a single 8x24GB GPU machine.
Democratization: Makes full fine-tuning (not just LoRA) accessible to researchers without multi-node GPU clusters.
Optimizer memory research: Joined the 2023 wave of optimizer memory innovations (8-bit Adam, AdaFactor, GaLore) democratizing large-model tuning.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack