🚀NEW LABGetting Started with Claude AgentsStart lab
Multimodal · Evaluation · Data

DataComp

First page
DataComp
Paper summary

A multimodal dataset benchmark with 12.8B image-text pairs.

Ask this paper

Key points
01

Scale and scope: 12.8 billion image-text pairs - one of the largest multimodal datasets ever released.

02

Benchmark framework: Provides a benchmark where researchers compete to find the best data subset, not just train the best model on fixed data.

03

Data-centric AI: Emphasizes data curation as the primary research axis, with model architecture and training held constant.

04

Data research infrastructure: Enabled a wave of data-filtering research (DataComp-XL, fastText filtering) that significantly advanced multimodal model training.

Every Monday
Get next week’s papers.
Subscribe on Substack