🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Reasoning · Reinforcement Learning · Training

ReFT (Reinforced Fine-Tuning)

Free while signed in. Answers cite the passages they came from.

First page
ReFT (Reinforced Fine-Tuning)
The curator’s take

ByteDance's ReFT enhances LLM reasoning by combining supervised fine-tuning with online RL that samples alternative reasoning paths, without a learned reward model.

Key points
01

Two-stage recipe: Standard SFT gives the model a reasonable starting policy; an online RL phase then refines it by sampling reasoning paths and rewarding those that yield correct final answers.

02

No reward model needed: Unlike RLHF, ReFT uses the ground-truth answer itself as the reward signal, avoiding the cost and instability of training a separate reward model.

03

Math problem-solving wins: Strong gains on GSM8K and MATH, with better out-of-distribution generalization than pure SFT at the same compute.

04

Simple conceptual lesson: The paper is an early demonstration that online RL from verifiable signals can outperform supervised recipes on reasoning - a theme that would dominate the subsequent year's research.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack