🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

SimPO

First page
SimPO
Paper summary

a simpler and more effective approach for preference optimization with a reference-free reward; uses the average log probability of a sequence as an implicit reward (i.e., no reference model required) which makes it more compute and memory efficient; demonstrates that it outperforms existing approaches like DPO and claims to produce the strongest 8B open-source model.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack