🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reinforcement Learning · Evaluation

Self-Rewarding Language Models

Free while signed in. Answers cite the passages they came from.

First page
Self-Rewarding Language Models
The curator’s take

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.

Key points
01

LLM-as-a-Judge inside training: The same model generates candidate responses *and* scores them using LLM-as-a-Judge prompting, producing preference pairs automatically.

02

Iterative DPO: Preference pairs are used in DPO-style instruction-following training, with each iteration producing both a stronger policy and a stronger judge.

03

Three iterations: Three rounds of self-rewarding fine-tuning on Llama 2 70B yield a model that outperforms Claude 2 and Gemini Pro on AlpacaEval 2.0.

04

Open-loop risk: Raises the concerning but productive question of whether models trained on their own evaluations eventually plateau, diverge, or keep improving - a live research area following this paper.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack