DataComp
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsA multimodal dataset benchmark with 12.8B image-text pairs.
01
Scale and scope: 12.8 billion image-text pairs - one of the largest multimodal datasets ever released.
02
Benchmark framework: Provides a benchmark where researchers compete to find the best data subset, not just train the best model on fixed data.
03
Data-centric AI: Emphasizes data curation as the primary research axis, with model architecture and training held constant.
04
Data research infrastructure: Enabled a wave of data-filtering research (DataComp-XL, fastText filtering) that significantly advanced multimodal model training.
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack