🚀NEW LABGetting Started with Claude AgentsStart lab
Efficiency

Sparse-Quantized Representation (SpQR)

First page
Sparse-Quantized Representation (SpQR)
Paper summary

Tim Dettmers' near-lossless LLM compression technique.

Ask this paper

Key points
01

4.75-bit inference: Enables LLM inference at 4.75 bits per parameter with a 15% speedup over FP16 baselines.

02

Near-lossless: Maintains model quality close to full-precision, with degradation measured in fractions of a percent on standard benchmarks.

03

Outlier-aware quantization: Identifies and preserves sensitive "outlier" weights in higher precision while aggressively quantizing the rest.

04

Quantization lineage: Part of Dettmers' influential quantization research (LLM.int8, QLoRA, SpQR) that made large-model inference accessible on consumer hardware.

Every Monday
Get next week’s papers.
Subscribe on Substack