Self-Play Fine-Tuning (SPIN)
Free while signed in. Answers cite the passages they came from.

SPIN shows that a supervised fine-tuned LLM can keep improving via self-play alone, without any additional human annotations.
Self-play loop: At each iteration, the current policy generates responses, and the model is then trained to distinguish its own responses from human-annotated ones - squeezing more signal out of the original SFT dataset.
No extra labels: Uses only the existing SFT dataset across iterations; no reward model and no new human annotations are required to drive further gains.
Beats DPO with GPT-4 labels: SPIN outperforms DPO training that uses GPT-4 preference labels on the same SFT data, a striking result given DPO's access to richer signal.
Scaling signal: Suggests that the SFT dataset itself contains more information than a single training pass extracts, with self-play acting as a lightweight amplifier.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack