BabyLLM Challenge Findings
Free while signed in. Answers cite the passages they came from.

Reports results from a challenge on sample-efficient pretraining using a developmentally plausible corpus.
Constrained pretraining: Participants pretrain on a small, child-directed-style corpus rather than on internet-scale data, testing how efficiently models can learn from limited input.
LTG BERT wins: The winning submission, LTG BERT, beat Llama 2 70B on 3 of 4 evaluations despite vastly less training data.
Data preprocessing pays: Strong-performing entries relied heavily on data preprocessing and training on shorter contexts, challenging assumptions about long-context training for small data.
Cognitive-science bridge: Provides an empirical platform connecting language-model training to developmental psycholinguistics, informing both fields.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack