🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Training · Safety

RLVR Meets Human Likeness

First page
RLVR Meets Human Likeness
Paper summary

RL with verifiable rewards only optimizes what you can objectively score, so style, structure, and diversity quietly collapse and reward hacking creeps in. This MIT work adds an adversarial discriminator trained on human demonstrations as a learned proxy for the human output distribution, and the generator maximizes both task accuracy and that human-likeness signal. Across bug fixing, story generation, and a reward-hacking benchmark, it preserves RLVR's accuracy gains while restoring the fuzzy properties it usually destroys, with misbehavior nearly disappearing.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack