🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 12, 2026
Evaluation

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers

First page
Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers
The curator’s take

Hanhua Hong, Chenghua Lin and colleagues (Manchester, IQuest Research, Beihang and Langboat) introduce AgentActionBench for the NLPCC 2026 shared task, which grades how agents reproduce paper experiments by recording their actions rather than only checking the final repository.

Ask this paper

Key points
01

Process evaluation: An MCP-based Action Recorder captures agent behavior during reproduction, and paper-specific rubrics score the recorded trace.

02

Coverage: 150 papers, 120 in machine learning and 30 in AI for science, extending beyond ML-only reproduction benchmarks.

03

Rubrics at scale: A human-annotated 10% subset validates model-assisted augmentation to more than 10,000 rubric items, with strong Pearson and Spearman agreement.

04

Finding: Current systems remain limited, and execution is the main bottleneck.

Abstract

Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus on final repositories and are typically limited to machine learning (ML). We introduce AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains. Our framework uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics. AgentActionBench contains 150 papers, including 120 ML papers and 30 AI4Science papers. A human-annotated subset covering 10% of the benchmark provides validation data, while model-assisted augmentation expands the full benchmark to more than 10,000 rubric items. Experimental results show that current systems remain limited, with execution as the primary bottleneck. Meanwhile, the strong Pearson and Spearman correlations between model-generated and human-annotated rubrics validate the reliability of our scalable rubric-generation approach.

Every Monday
Get next week’s papers.
Subscribe on Substack