🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

WARM (Weighted Averaged Reward Models)

First page
WARM (Weighted Averaged Reward Models)
Paper summary

WARM averages multiple fine-tuned reward models in weight space rather than ensembling their predictions, dramatically reducing RLHF inference cost.

Ask this paper

Key points
01

Weight-space averaging: Trains several reward models from the same pretrained initialization and averages their weights into a single model, exploiting the "linear mode connectivity" observed in fine-tuned LLMs.

02

Efficient vs. prediction ensembling: Gives most of the robustness gains of prediction ensembling while running at the cost of a single reward model at inference time.

03

Improved alignment: Policies trained against WARM produce higher-quality and more aligned generations than policies trained against any single reward model.

04

Robust to reward hacking: Because WARM smooths over idiosyncratic flaws of individual reward models, it is measurably more resistant to reward hacking during PPO-style optimization.

Every Monday
Get next week’s papers.
Subscribe on Substack