TinyStories
First page

Paper summary
Explores how small LMs can be and still speak coherent English.
Ask this paper
01
Synthetic story dataset: Creates a dataset of short stories using words understandable to 3-4 year olds, generated by GPT-3.5/GPT-4.
02
Tiny but fluent: Shows that very small models (1-10M parameters) trained on this focused data can produce coherent multi-paragraph stories.
03
Reasoning emergence: Even tiny models demonstrate reasoning and instruction-following capabilities when trained on the right data.
04
Data-quality evidence: A foundational piece in the argument that data quality beats scale for many capabilities, influencing Phi series and later SLM work.