When Better Turns Do Not Make Better Agents: Diagnosing the Gap Between Next-Turn Metrics and Workflow Success

Laskar, Fu and colleagues (Dialpad) test whether improving next-turn metrics under gold history predicts better autonomous multi-turn workflow execution, and find that it does not.
Ask this paper
Setup. Pre-SFT and SFT Qwen3 (4B, 14B) and Gemma 3 (4B, 12B) models are evaluated on multi-turn customer-support workflows.
Next-turn gains. SFT improves text-turn success, and overall next-turn success rises for every model under gold-history evaluation.
No transfer. None of the four SFT models succeeds under holistic workflow evaluation; strict trajectory completion peaks at 10.4%.
Recommendation. Report text quality, local action correctness, tool execution and end-to-end completion separately. Accepted to the REALM workshop at EMNLP 2026.
Abstract
Agent models are frequently evaluated one decision at a time, where the model predicts the next action based on the gold interaction history, which is scored against a reference. We investigate whether improvement under this protocol is predictive of improved autonomous workflow execution. We study pre-SFT and supervised fine-tuned (SFT) Qwen3 models at 4B and 14B parameters and Gemma 3 models at 4B and 12B parameters on multi-turn customer-support workflows. We find that SFT consistently improves text-turn success, and that overall next-turn success increases for every model under gold-history evaluation. However, these improvements do not transfer to autonomous workflow execution. Tool-specific gains also vary across metrics and models. None of the four SFT models succeeds under holistic workflow evaluation, with strict trajectory completion reaching at most 10.4% workflow success. Our results show that next-turn evaluation is not a reliable proxy for workflow success, motivating separate reporting of text quality, local action correctness, tool execution, and end-to-end task completion.