🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 23, 2026
Evaluation · Agents

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

First page
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
The curator’s take

An, Jang, Kim, Lee, Park and Lee (KRAFTON) release AgentVidBench, a multi-hop video QA benchmark that tests spatial, temporal and causal reasoning in multimodal agents and scores the solution trajectory as well as the final answer.

Ask this paper

Key points
01

Step traces. Each question comes with a step-by-step solution trace, so evaluation checks whether the agent actually gathered the evidence that justifies its answer.

02

Single-turn weakness. Across 12 proprietary and open MLLMs, single-turn performance is limited.

03

Agentic gain. Placing the same models inside current agentic workflows generally improves both accuracy and trajectory scores.

04

Baseline and release. The authors add a simple agentic strategy as a competitive baseline and release code and data on GitHub and Hugging Face.

Abstract

Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these agentic capabilities. To bridge this gap, we introduce AgentVidBench, a multi-hop video question answering benchmark focused on evaluating the spatial, temporal, and causal reasoning capabilities of MLLM agents. Beyond standard question-answer pairs, AgentVidBench provides step-by-step solution traces to support trajectory evaluation that assesses whether agents explicitly acquire the evidence needed to justify their answers. Experiments with 12 proprietary and open-source MLLMs show that single-turn performance remains limited on AgentVidBench, while integrating these models into state-of-the-art agentic workflows generally improves performance with respect to both accuracy and trajectory scores. We further present a simple yet effective agentic strategy that serves as a competitive baseline on AgentVidBench, establishing our benchmark as a holistic testbed for future research on agentic video understanding. Code and datasets are available at https://github.com/krafton-ai/agentvidbench and https://huggingface.co/datasets/agentvidbench/agentvidbench.

Every Monday
Get next week’s papers.
Subscribe on Substack