🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Data · Reasoning · Training

A Primer on Post-Training Reasoning Data

Free while signed in. Answers cite the passages they came from.

First page
A Primer on Post-Training Reasoning Data
The curator’s take

This primer is the first to pull the scattered post-training reasoning-data literature into one place, synthesizing over 150 public studies and system reports that previously lived across dataset papers, RL write-ups, and lab reports. It organizes the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. The key reframing is that a reasoning-data item is more than a prompt-response pair: it packages a problem or state, model behavior, judging feedback, and attribution metadata, with usefulness defined relative to the verifier and the rest of the corpus rather than in isolation.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack