Self-Rewarding Language Models

Meta shows that an LLM can act as both actor and judge in its own alignment loop, generating training data without any external reward model.
Ask this paper
LLM-as-a-Judge inside training: The same model generates candidate responses *and* scores them using LLM-as-a-Judge prompting, producing preference pairs automatically.
Iterative DPO: Preference pairs are used in DPO-style instruction-following training, with each iteration producing both a stronger policy and a stronger judge.
Three iterations: Three rounds of self-rewarding fine-tuning on Llama 2 70B yield a model that outperforms Claude 2 and Gemini Pro on AlpacaEval 2.0.
Open-loop risk: Raises the concerning but productive question of whether models trained on their own evaluations eventually plateau, diverge, or keep improving - a live research area following this paper.