🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Efficiency

QLoRA

First page
QLoRA
Paper summary

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.

Ask this paper

Key points
01

4-bit NF4 quantization: Introduces the NormalFloat 4-bit datatype optimized for normally-distributed weights with double-quantization for further memory savings.

02

Paged optimizers: Uses paged NVIDIA Unified Memory to handle optimizer state memory spikes without OOM failures.

03

16-bit quality: Achieves quality matching full 16-bit fine-tuning despite aggressive quantization during training.

04

Community fine-tuning enabler: Arguably the single most impactful 2023 paper for democratizing LLM fine-tuning - powered thousands of community checkpoints on Hugging Face.

Every Monday
Get next week’s papers.
Subscribe on Substack