WARM (Weighted Averaged Reward Models)
Free while signed in. Answers cite the passages they came from.

WARM averages multiple fine-tuned reward models in weight space rather than ensembling their predictions, dramatically reducing RLHF inference cost.
Weight-space averaging: Trains several reward models from the same pretrained initialization and averages their weights into a single model, exploiting the "linear mode connectivity" observed in fine-tuned LLMs.
Efficient vs. prediction ensembling: Gives most of the robustness gains of prediction ensembling while running at the cost of a single reward model at inference time.
Improved alignment: Policies trained against WARM produce higher-quality and more aligned generations than policies trained against any single reward model.
Robust to reward hacking: Because WARM smooths over idiosyncratic flaws of individual reward models, it is measurably more resistant to reward hacking during PPO-style optimization.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack