Rephrasing the Web (WRAP)
Free while signed in. Answers cite the passages they came from.

WRAP uses an off-the-shelf instruction-tuned model to paraphrase web documents into styles like "Wikipedia" or "question-answer format" and trains on the mixture of real + synthetic rephrases.
Style-conditioned rephrasing: A frozen instruction-tuned model rewrites the same content into multiple stylistic variants, expanding effective diversity without changing meaning.
~3x faster pretraining: Training on real + rephrased web data reaches the same perplexity as a real-only baseline in roughly 3x fewer steps on the Pile.
Downstream gains: Beyond perplexity, zero-shot QA accuracy improves across 13 tasks compared to the real-only baseline at matched compute.
Data quality lever: Suggests that converting messy web text into cleaner stylistic forms at scale is a simple and effective lever for pretraining efficiency, without needing new human data.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack