Best Practices and Lessons on Synthetic Data

Google DeepMind's survey-style position paper on synthetic data for LLMs. It covers applications, quality-assurance principles, and the open challenges of factuality, fidelity, bias, and privacy.
Ask this paper
Why synthetic data now: Addresses the shortage of large, diverse, high-quality natural data and the privacy constraints around using real user content, positioning synthetic data as a complementary source rather than a replacement.
Quality-assurance triad: Emphasizes factuality (claims are true), fidelity (distribution matches real use), and unbiasedness (no over/under-representation) as the three tests every synthetic pipeline should run.
Responsible use: Discusses provenance tagging, contamination risks with eval sets, and how to avoid model-collapse feedback loops when synthetic data is used for pretraining.
Open directions: Calls out evaluation of synthetic pipelines, hybrid natural+synthetic recipes, and privacy-preserving generation (e.g., differential privacy) as the main frontiers.