Datasets for LLMs: A Comprehensive Survey

A 180+-page survey that catalogs and analyzes the datasets that underpin modern LLM training and evaluation.
Ask this paper
Five-category taxonomy: Organizes the space into pretraining corpora, instruction fine-tuning datasets, preference datasets, evaluation datasets, and traditional NLP datasets.
Scale of the review: Covers 444 datasets spanning eight language families and 32 domains, including 774.5 TB of pretraining data and 700M+ instances across non-pretraining dataset types.
Twenty-dimension statistics: Each dataset is summarized across 20 attributes (license, language, size, task type, etc.), enabling systematic comparison rather than anecdotal selection.
Challenges and gaps: Identifies open issues in dataset construction (licensing, provenance, bias, contamination) and outlines future research directions for responsible LLM data pipelines.