🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 22, 2026
Agents · Evaluation

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

First page
Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
The curator’s take

Xinshuai Guo, Junjie Wu and colleagues at Tencent Hunyuan and Tsinghua University propose DualViewEval, which compresses expensive agent benchmarks into small task subsets by modeling both final scores and process signals from agent trajectories.

Ask this paper

Key points
01

Process signals. An analysis of large-scale trajectories finds six process signals that correlate with final agent performance, which outcome-only compression methods ignore.

02

Dual-view selection. The method learns an exact-size miniset from outcome relations and process relations jointly, then predicts full-benchmark scores from it.

03

Compression ratio. With 20 tasks it compresses APEX-Agents and BFCL by 24x to 40x and cuts mean absolute error by 14.5% to 28.2% against the strongest baselines.

04

Ranking fidelity. On SWE-bench Verified it improves Kendall's tau by up to 7.2% relative to EssenceBench, and it wins on all five benchmarks tested.

05

Diagnostic minisets. The selected tasks also separate agents by capability, so the miniset doubles as a compact diagnostic during model development.

Abstract

Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are systematically associated with final agent performance. To disentangle agent performance redundancy from a complete perspective, we propose DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores. Across five agent benchmarks and five representative baselines, DualViewEval achieves the best results in all datasets. With only 20 tasks, it achieves $24\times$--$40\times$ compression on APEX-Agents and BFCL, reducing mean absolute error (MAE) by $14.5\%$--$28.2\%$ over the strongest competitors while improving Kendall's $τ$ by up to $7.2\%$ relative to EssenceBench on SWE-bench Verified. The selected minisets further reveal capability differences among different agents, providing compact and diagnostic feedback for efficient agentic model development.

Every Monday
Get next week’s papers.
Subscribe on Substack