🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 5, 2026
Evaluation · Agents

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training

First page
UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
The curator’s take

Ashish Jain (Sarvam AI) and Armaan Sandhu (UMass Amherst) introduce UserProxyBench, which scores the simulated user in tau-bench-style agent benchmarks on whether it followed its private instructions, separately from whether the agent succeeded.

Ask this paper

Key points
01

Metric. The User Fidelity Score checks adherence to the benchmark's own user specification with task-grounded rubric criteria, judged independently of agent reward.

02

The user changes the score. Holding the agent fixed at GPT-5.5 and swapping only the user proxy across 375 enterprise tasks moves mean task reward by 15.2 points.

03

Hidden violations. 24.4% of successful episodes contain a user-specification violation, so task reward does not reveal whether the user behaved correctly.

04

Premature disclosure. The most common failure is the user volunteering information before the agent asks. It barely affects reward, but in successful episodes the agent makes 1.06 fewer tool calls, so the benchmark is measuring a different interaction than intended.

05

Cost-fidelity frontier. Across seven proxies the authors chart cost against fidelity so teams can pick the cheapest simulator that meets a required fidelity level.

Abstract

Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding the agent fixed at GPT-5.5 and varying only the user proxy across 375 enterprise tasks changes mean task reward by 15.2 points, while 24.4% of successful episodes contain a user-specification violation. The dominant failure is premature disclosure: users provide information before it is requested. This behavior has little effect on task reward, yet among successful episodes it causes the agent to make 1.06 fewer tool calls on average, changing the interaction being evaluated while preserving the reward. Finally, across seven proxies we identify an empirical cost-fidelity frontier, enabling practitioners to select the least expensive simulator that satisfies a required fidelity level.

Every Monday
Get next week’s papers.
Subscribe on Substack