🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Data

DoReMi

First page
DoReMi
Paper summary

Optimizes data mixtures for faster language model pretraining.

Ask this paper

Key points
01

Proxy-model reweighting: Trains a small 280M proxy model with group-DRO to derive optimal domain weights for the actual pretraining mixture.

02

Scale transfer: Weights found by 280M proxy transfer to training 8B models (30x larger) without retuning.

03

Training speedup: Achieves faster convergence and better downstream performance than uniform or human-tuned mixtures.

04

Data mixture research: Kicked off a wave of data-mixture optimization work that became central to 2024 pretraining recipes (Llama 3, DCLM).

Every Monday
Get next week’s papers.
Subscribe on Substack