🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Agents

Training Advisors for LLM Agents from Task Outcomes

First page
Training Advisors for LLM Agents from Task Outcomes
The curator’s take

Sergei Polezhaev, Barys Liskavets, Ori Press and Alexander Golubev (Nebius AI) introduce Caddie, which trains a critic model to give natural-language advice to a frozen agent mid-task, using only whether the agent eventually succeeds as reward.

Ask this paper

Key points
01

Training signal. No step-level labels or reference critiques. The critic is optimized with RL on the terminal outcome of the frozen base model's continuation after receiving the critique.

02

Main result. A Qwen3-4B critic trained on multi-hop QA with one base model raises Qwen3-4B's MuSiQue success rate by more than 25 points, above Kimi K3 running without a critic.

03

Transfer across models. The same critic improves four base models of different sizes and architectures, three of which were never seen during critic training.

04

Transfer across tasks. Without further training it also helps on out-of-domain interactive benchmarks, tau3 and DeepDive.

05

Invocation. Agents can decide at inference time when to ask the critic for help.

Abstract

Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $\tau^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.

Every Monday
Get next week’s papers.
Subscribe on Substack