🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Reinforcement Learning

Hybrid Reinforcement

First page
Hybrid Reinforcement
Paper summary

HERO (Hybrid Ensemble Reward Optimization) is a reinforcement learning framework that combines binary verifier feedback with continuous reward-model signals to improve LLM reasoning. By using stratified normalization and variance-aware weighting, HERO balances correctness and nuance, outperforming verifier-only and RM-only methods on diverse math reasoning benchmarks and enhancing performance on both verifiable and ambiguous tasks.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack