QLoRA
Free while signed in. Answers cite the passages they came from.

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.
4-bit NF4 quantization: Introduces the NormalFloat 4-bit datatype optimized for normally-distributed weights with double-quantization for further memory savings.
Paged optimizers: Uses paged NVIDIA Unified Memory to handle optimizer state memory spikes without OOM failures.
16-bit quality: Achieves quality matching full 16-bit fine-tuning despite aggressive quantization during training.
Community fine-tuning enabler: Arguably the single most impactful 2023 paper for democratizing LLM fine-tuning - powered thousands of community checkpoints on Hugging Face.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack