Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

Yu Li and colleagues at Southeast University introduce CITA, which trains a Comparative Inference Model to rank candidate next tool calls by their expected effect on final task success, and uses it to guide GRPO training and inference; the paper is a NeurIPS 2026 poster.
Ask this paper
Pilot analysis. 95% of failed tool-use trajectories stay unrecoverable after one backtracking step and 76% after four. At 49% of fork points, the tool the model prefers has a lower empirical success rate than an alternative in the same context.
Comparative value. Logged trajectories only show the chosen call, so CIM learns from paired evidence built from real trajectories, a Bayesian tool-graph simulator and LLM comparisons, and estimates relative long-horizon value between candidate calls.
Results. On Toolathlon, TOUCAN and TRAJECT-Bench with Qwen2.5-7B and Llama3.1-8B backbones, CITA is best in all six settings, with average gains of 7.55 Tool F1 points and 9.93 task-success points over the strongest prior method.
Scope. Experiments assume logged trajectories and a known tool graph; settings where tools are added or removed over time are left to future work.
Abstract
Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.