🚀NEW LABGetting Started with Claude AgentsStart lab
Data

Rephrasing the Web (WRAP)

First page
Rephrasing the Web (WRAP)
Paper summary

WRAP uses an off-the-shelf instruction-tuned model to paraphrase web documents into styles like "Wikipedia" or "question-answer format" and trains on the mixture of real + synthetic rephrases.

Ask this paper

Key points
01

Style-conditioned rephrasing: A frozen instruction-tuned model rewrites the same content into multiple stylistic variants, expanding effective diversity without changing meaning.

02

~3x faster pretraining: Training on real + rephrased web data reaches the same perplexity as a real-only baseline in roughly 3x fewer steps on the Pile.

03

Downstream gains: Beyond perplexity, zero-shot QA accuracy improves across 13 tasks compared to the real-only baseline at matched compute.

04

Data quality lever: Suggests that converting messy web text into cleaner stylistic forms at scale is a simple and effective lever for pretraining efficiency, without needing new human data.

Every Monday
Get next week’s papers.
Subscribe on Substack