TinyStories
Free while signed in. Answers cite the passages they came from.

Explores how small LMs can be and still speak coherent English.
Synthetic story dataset: Creates a dataset of short stories using words understandable to 3-4 year olds, generated by GPT-3.5/GPT-4.
Tiny but fluent: Shows that very small models (1-10M parameters) trained on this focused data can produce coherent multi-paragraph stories.
Reasoning emergence: Even tiny models demonstrate reasoning and instruction-following capabilities when trained on the right data.
Data-quality evidence: A foundational piece in the argument that data quality beats scale for many capabilities, influencing Phi series and later SLM work.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack