🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

Process Reinforcement through Implicit Rewards

Paper preview
Process Reinforcement through Implicit Rewards
Paper summary

a framework for online reinforcement learning that uses process rewards to improve language model reasoning; the proposed algorithm combines online prompt filtering, RLOO return/advantage estimation, PPO loss, and implicit process reward modeling online updates; on their model, Eurus-2-7B-PRIME, achieves 26.7% pass@1 on AIME 2024, surpassing GPT-4 and other models, using only 1/10 of the training data compared to similar models.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack