🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Multimodal · Training

Beyond Scalar Rewards

First page
Beyond Scalar Rewards
Paper summary

Reward models usually compress a judgment into a single scalar, but this paper argues human preferences are better captured as score distributions, and proposes Z-Reward, which internalizes reasoning into a predicted distribution before scoring. A large vision-language teacher does the reasoning-heavy judgment and is distilled into a compact student for efficient deployment, with the 27B teacher reaching 89.6% human-preference accuracy and the 9B student nearly matching it at 88.6%. Used as a reinforcement learning signal, it delivers a 41.3% net preference improvement over a supervised baseline, beating GRPO and other reward methods.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack