🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Evaluation

Training LLMs to Self-Correct via RL

First page
Training LLMs to Self-Correct via RL
Paper summary

develops a multi-turn online reinforcement learning to improve the capabilities of an LLM to self-correct; it’s based entirely on self-generated data; SFT is shown to be ineffective at learning self-correction and suffers from distribution mismatch between training data and model responses; proposes a two-stage approach that first optimizes correction behavior and then uses a reward bonus to amplify self-correction during training; when applied to Gemini 1.0 Pro and 1.5 Flash models, it achieves state-of-the-art self-correction performance, improving the base models’ self-correction by 15.6% and 9.1% respectively on the MATH and HumanEval benchmarks.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack