Sparse-Quantized Representation (SpQR)
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsTim Dettmers' near-lossless LLM compression technique.
01
4.75-bit inference: Enables LLM inference at 4.75 bits per parameter with a 15% speedup over FP16 baselines.
02
Near-lossless: Maintains model quality close to full-precision, with degradation measured in fractions of a percent on standard benchmarks.
03
Outlier-aware quantization: Identifies and preserves sensitive "outlier" weights in higher precision while aggressively quantizing the rest.
04
Quantization lineage: Part of Dettmers' influential quantization research (LLM.int8, QLoRA, SpQR) that made large-model inference accessible on consumer hardware.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack