🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 23, 2026
Agents · Evaluation

When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

First page
When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success
The curator’s take

Laskar, Fu and colleagues (Dialpad) test whether improving next-turn metrics under gold history predicts better autonomous multi-turn workflow execution, and find that it does not.

Ask this paper

Key points
01

Setup. Pre-SFT and SFT Qwen3 (4B, 14B) and Gemma 3 (4B, 12B) models are evaluated on multi-turn customer-support workflows.

02

Next-turn gains. SFT improves text-turn success, and overall next-turn success rises for every model under gold-history evaluation.

03

No transfer. None of the four SFT models succeeds under holistic workflow evaluation; strict trajectory completion peaks at 10.4%.

04

Recommendation. Report text quality, local action correctness, tool execution and end-to-end completion separately. Accepted to the REALM workshop at EMNLP 2026.

Abstract

Agent models are frequently evaluated one decision at a time, where the model predicts the next action based on the gold interaction history, which is scored against a reference. We investigate whether improvement under this protocol is predictive of improved autonomous workflow execution. We study pre-SFT and supervised fine-tuned (SFT) Qwen3 models at 4B and 14B parameters and Gemma 3 models at 4B and 12B parameters on multi-turn customer-support workflows. We find that SFT consistently improves text-turn success, and that overall next-turn success increases for every model under gold-history evaluation. However, these improvements do not transfer to autonomous workflow execution. Tool-specific gains also vary across metrics and models. None of the four SFT models succeeds under holistic workflow evaluation, with strict trajectory completion reaching at most 10.4% workflow success. Our results show that next-turn evaluation is not a reliable proxy for workflow success, motivating separate reporting of text quality, local action correctness, tool execution, and end-to-end task completion.

Every Monday
Get next week’s papers.
Subscribe on Substack