Fine-Tuning Language Models with Just Forward Passes (MeZO)
First page

Paper summary
A memory-efficient zeroth-order optimizer for LLM fine-tuning.
Ask this paper
01
No backpropagation: Uses a memory-efficient zeroth-order SGD algorithm that requires only forward passes, eliminating the memory overhead of backprop.
02
Inference-like memory: Fine-tunes large LLMs with the same memory footprint as inference - democratizes full-parameter fine-tuning.
03
Comparable quality: Reaches comparable quality to backpropagation-based fine-tuning on many tasks despite using only forward passes.
04
Memory-constrained tuning: Opens new possibilities for fine-tuning huge models on modest hardware by trading compute for memory.