QLoRA
First page

Paper summary
Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.
Ask this paper
01
4-bit NF4 quantization: Introduces the NormalFloat 4-bit datatype optimized for normally-distributed weights with double-quantization for further memory savings.
02
Paged optimizers: Uses paged NVIDIA Unified Memory to handle optimizer state memory spikes without OOM failures.
03
16-bit quality: Achieves quality matching full 16-bit fine-tuning despite aggressive quantization during training.
04
Community fine-tuning enabler: Arguably the single most impactful 2023 paper for democratizing LLM fine-tuning - powered thousands of community checkpoints on Hugging Face.