DataComp
First page

Paper summary
A multimodal dataset benchmark with 12.8B image-text pairs.
Ask this paper
01
Scale and scope: 12.8 billion image-text pairs - one of the largest multimodal datasets ever released.
02
Benchmark framework: Provides a benchmark where researchers compete to find the best data subset, not just train the best model on fixed data.
03
Data-centric AI: Emphasizes data curation as the primary research axis, with model architecture and training held constant.
04
Data research infrastructure: Enabled a wave of data-filtering research (DataComp-XL, fastText filtering) that significantly advanced multimodal model training.