🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Efficiency

Sparse-Quantized Representation (SpQR)

Free while signed in. Answers cite the passages they came from.

First page
Sparse-Quantized Representation (SpQR)
The curator’s take

Tim Dettmers' near-lossless LLM compression technique.

Key points
01

4.75-bit inference: Enables LLM inference at 4.75 bits per parameter with a 15% speedup over FP16 baselines.

02

Near-lossless: Maintains model quality close to full-precision, with degradation measured in fractions of a percent on standard benchmarks.

03

Outlier-aware quantization: Identifies and preserves sensitive "outlier" weights in higher precision while aggressively quantizing the rest.

04

Quantization lineage: Part of Dettmers' influential quantization research (LLM.int8, QLoRA, SpQR) that made large-model inference accessible on consumer hardware.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack