DoReMi
Free while signed in. Answers cite the passages they came from.
First page

The curator’s take
Key pointsOptimizes data mixtures for faster language model pretraining.
01
Proxy-model reweighting: Trains a small 280M proxy model with group-DRO to derive optimal domain weights for the actual pretraining mixture.
02
Scale transfer: Weights found by 280M proxy transfer to training 8B models (30x larger) without retuning.
03
Training speedup: Achieves faster convergence and better downstream performance than uniform or human-tuned mixtures.
04
Data mixture research: Kicked off a wave of data-mixture optimization work that became central to 2024 pretraining recipes (Llama 3, DCLM).
Every Monday
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack