Data Management for LLMs
Free while signed in. Answers cite the passages they came from.

A survey of data-management research for LLM pretraining and supervised fine-tuning stages.
Pretraining data: Covers data quantity, quality filtering, deduplication, domain composition, and curriculum strategies for large-scale pretraining.
SFT data: Reviews instruction-data generation, quality filtering, diversity metrics, and the emerging literature on "less is more" for SFT.
Domain and task composition: Examines how task mixing affects generalization vs. specialization in fine-tuning.
Open challenges: Identifies dataset contamination, deduplication at trillion-token scale, and reproducible data recipes as the top open problems.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack