🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training · Data

Aligning LLMs to Quote from Pre-Training Data (Quote-Tuning)

Free while signed in. Answers cite the passages they came from.

First page
Aligning LLMs to Quote from Pre-Training Data (Quote-Tuning)
The curator’s take

Quote-Tuning aligns LLMs to quote verbatim from trusted pre-training sources, turning the attribution step from post-hoc fact-checking into a built-in model behavior.

Key points
01

Membership inference at train time: A fast membership-inference function checks whether generated spans exist verbatim in a trusted corpus, producing a reward signal without any human annotation.

02

Preference-based alignment: The authors build a synthetic preference dataset (quoted vs non-quoted outputs) and align the model with preference optimization, teaching it when to quote.

03

Strong verbatim gains: Quote-Tuning achieves up to a 130% relative increase in verbatim quotes from high-quality documents while preserving response quality across tasks, domains, and model families.

04

Verification advantage: Because quoted passages can be matched exactly to the source, downstream verification becomes trivial, helping regulated domains like medicine, law, and journalism.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack