🚀NEW LABGetting Started with Claude AgentsStart lab
Data

TinyStories

First page
TinyStories
Paper summary

Explores how small LMs can be and still speak coherent English.

Ask this paper

Key points
01

Synthetic story dataset: Creates a dataset of short stories using words understandable to 3-4 year olds, generated by GPT-3.5/GPT-4.

02

Tiny but fluent: Shows that very small models (1-10M parameters) trained on this focused data can produce coherent multi-paragraph stories.

03

Reasoning emergence: Even tiny models demonstrate reasoning and instruction-following capabilities when trained on the right data.

04

Data-quality evidence: A foundational piece in the argument that data quality beats scale for many capabilities, influencing Phi series and later SLM work.

Every Monday
Get next week’s papers.
Subscribe on Substack