🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 13, 2026
Multimodal · Evaluation · Agents

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

First page
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
The curator’s take

Jiaqiang Li, Tao Gui and colleagues (Fudan NLP Group with Shanghai AI Laboratory) build Sci-MMR, a benchmark that checks whether multimodal research agents recover the full evidence chain behind a scientific answer, and find answer accuracy runs more than 20 points ahead of evidence recovery.

Ask this paper

Key points
01

Benchmark: 235 multi-hop tasks across four disciplines, built on argument graphs linking claims, citations, visual evidence and figure regions, with about nine figure panels per task.

02

Answer versus evidence: Across eight frontier multimodal models, answer accuracy exceeds complete-evidence recovery by more than 20%.

03

Acquisition failures: 57.2% of failures come from extracting evidence from figures. Cropping tools add 4.5 points while gold evidence adds up to 37.0.

04

Integration failures: 31.8% of failures come from turning available evidence into conclusions. Even with gold evidence the best model reaches 69.1% on the hardest tasks.

Abstract

Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents

Every Monday
Get next week’s papers.
Subscribe on Substack