🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Evaluation

Self-Taught Evaluators

First page
Self-Taught Evaluators
Paper summary

an approach to improve model-based evaluators using synthetic training data only; it first generates contrasting outputs (good and bad model responses) and trains an LLM-as-a-Judge to produce reasoning traces and final judgments; the self-improvement scheme repeats the training process in an iterative way using its improved predictions; claims to outperform LLM-judges such as GPT-4 and match top-performing reward models trained on labeled examples; improves a strong LLM (Llama3-70BInstruct) from 75.4 to 88.3 (88.7 with majority vote) on RewardBench.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack