DoReMi
First page

Paper summary
Optimizes data mixtures for faster language model pretraining.
Ask this paper
01
Proxy-model reweighting: Trains a small 280M proxy model with group-DRO to derive optimal domain weights for the actual pretraining mixture.
02
Scale transfer: Weights found by 280M proxy transfer to training 8B models (30x larger) without retuning.
03
Training speedup: Achieves faster convergence and better downstream performance than uniform or human-tuned mixtures.
04
Data mixture research: Kicked off a wave of data-mixture optimization work that became central to 2024 pretraining recipes (Llama 3, DCLM).