🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

Self-Play Preference Optimization

First page
Self-Play Preference Optimization
Paper summary

proposes a self-play-based method for aligning language models; this optimation procedure treats the problem as a constant-sum two-player game to identify the Nash equilibrium policy; it addresses the shortcomings of DPO and IPO and effectively increases the log-likelihood of chose responses and decreases the rejected ones; SPPO outperforms DPO and IPO on MT-Bench and the Open LLM Leaderboard.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack