🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

Discovering Preference Optimization Algorithms with LLMs

First page
Discovering Preference Optimization Algorithms with LLMs
Paper summary

proposes LLM-driven objective discovery of state-of-the-art preference optimization; no human intervention is used and an LLM is prompted to propose and implement the preference optimization loss functions based on previously evaluated performance metrics; discovers an algorithm that adaptively combined logistic and exponential losses.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack