🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training · Efficiency

QLoRA

Free while signed in. Answers cite the passages they came from.

First page
QLoRA
The curator’s take

Tim Dettmers' breakthrough technique enabling 65B LLM fine-tuning on a single 48GB GPU.

Key points
01

4-bit NF4 quantization: Introduces the NormalFloat 4-bit datatype optimized for normally-distributed weights with double-quantization for further memory savings.

02

Paged optimizers: Uses paged NVIDIA Unified Memory to handle optimizer state memory spikes without OOM failures.

03

16-bit quality: Achieves quality matching full 16-bit fine-tuning despite aggressive quantization during training.

04

Community fine-tuning enabler: Arguably the single most impactful 2023 paper for democratizing LLM fine-tuning - powered thousands of community checkpoints on Hugging Face.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack