🚀NEW LABGetting Started with Claude AgentsStart lab
Agents · Reasoning · Training

ReST Meets ReAct

First page
ReST Meets ReAct
Paper summary

Proposes a ReAct-style agent that improves itself via reinforced self-training on its own reasoning traces.

Ask this paper

Key points
01

Self-critique ReAct: A ReAct-style agent with a self-critique step that evaluates its own reasoning and answers, generating a filterable trace dataset.

02

ReST-style iterative RL: Uses growing-batch RL from AI feedback to iteratively fine-tune on the agent's successful reasoning traces, improving over rounds without human labels.

03

Human-label-free: Minimizes human involvement; synthetic data with self-improvement from AI feedback is the primary training signal throughout.

04

Distillation to small models: The improved agent can be distilled into models 1-2 orders of magnitude smaller with comparable performance, dramatically cutting inference cost.

Every Monday
Get next week’s papers.
Subscribe on Substack