Best Practices and Lessons on Synthetic Data
Free while signed in. Answers cite the passages they came from.

Google DeepMind's survey-style position paper on synthetic data for LLMs. It covers applications, quality-assurance principles, and the open challenges of factuality, fidelity, bias, and privacy.
Why synthetic data now: Addresses the shortage of large, diverse, high-quality natural data and the privacy constraints around using real user content, positioning synthetic data as a complementary source rather than a replacement.
Quality-assurance triad: Emphasizes factuality (claims are true), fidelity (distribution matches real use), and unbiasedness (no over/under-representation) as the three tests every synthetic pipeline should run.
Responsible use: Discusses provenance tagging, contamination risks with eval sets, and how to avoid model-collapse feedback loops when synthetic data is used for pretraining.
Open directions: Calls out evaluation of synthetic pipelines, hybrid natural+synthetic recipes, and privacy-preserving generation (e.g., differential privacy) as the main frontiers.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack