🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Data

BabyLLM Challenge Findings

Paper preview
BabyLLM Challenge Findings
Paper summary

Reports results from a challenge on sample-efficient pretraining using a developmentally plausible corpus.

Ask this paper

Key points
01

Constrained pretraining: Participants pretrain on a small, child-directed-style corpus rather than on internet-scale data, testing how efficiently models can learn from limited input.

02

LTG BERT wins: The winning submission, LTG BERT, beat Llama 2 70B on 3 of 4 evaluations despite vastly less training data.

03

Data preprocessing pays: Strong-performing entries relied heavily on data preprocessing and training on shorter contexts, challenging assumptions about long-context training for small data.

04

Cognitive-science bridge: Provides an empirical platform connecting language-model training to developmental psycholinguistics, informing both fields.

Every Monday
Get next week’s papers.
Subscribe on Substack