🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Agents · Evaluation

ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

First page
ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
The curator’s take

Arkajyoti Chakraborty, Andreas Stolcke and colleagues at Uniphore (with UIUC) present ToolRACER, a pipeline that coordinates user, assistant and tool emulator models to generate validated multi-turn tool-calling conversations, many with non-cooperative users.

Ask this paper

Key points
01

Pipeline. Scenario generation, four judge-LLM evaluators that check syntax, dialogue consistency and faithfulness, and a refinement loop that repairs failed conversations using the evaluators' reasons.

02

Benchmark. ToolRACERBench covers six domains and 55 personas with 5.6K validated trajectories; about 66% contain failure-prone scenarios such as unhappy paths, impossible requests and adversarial users.

03

Training gains. Fine-tuning on ToolRACERBench alone raises the tau2-bench retail and airline macro average to 39.20%, above the base model and an APIGen-MT trained checkpoint.

04

Mixing. Combining the robust trajectories with in-domain data gives the largest gains for small models, with improvements on ACEBench end-to-end accuracy as well.

Abstract

Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66\% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as $\tau^2$-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across $\tau^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.

Every Monday
Get next week’s papers.
Subscribe on Substack