🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 28, 2026
Evaluation · Agents

BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents

First page
BabelArena: A Large-Scale Multilingual Benchmark for LLM Agents
The curator’s take

Peng Kuang, Minghao Wu and colleagues at Alibaba Token Hub (with UIUC, Northeastern and Monash) introduce BabelArena, a benchmark that ports existing English agent benchmarks into 23 languages while keeping tasks and graders executable.

Ask this paper

Key points
01

BabelFlow pipeline. An agentic workflow analyzes each benchmark's runtime dependencies, translates while preserving structure, and applies multi-layer verification plus human review. Against independent translation it cuts broken references from 20.0% to 2.5% on a 40-task audit.

02

Scale. 16,146 instances from 702 canonical tasks across four benchmark families, 13 domains and 23 languages.

03

No single winner. Among five frontier models none leads across all benchmark families, and language consistency drops 3.8 to 14.8 points outside English.

04

Low-resource failures differ. Lower-resource languages show larger shares of tool-use and control-flow errors, not just worse answers, and use up to about twice the English input tokens without longer interactions.

05

Drift toward English. English accounts for 91.2% of annotated language switches in sampled inconsistent trajectories, concentrated in structured-output tasks.

Abstract

Large language model (LLM) agents increasingly execute multi-step workflows through tool use and interaction with users and environments. However, current agent evaluations are largely English-centric, limiting our understanding of agent capabilities in multilingual settings. We introduce BabelFlow, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics. Using BabelFlow, we construct BabelArena, a task-aligned benchmark comprising 16,146 instances derived from 702 canonical tasks across four benchmark families, 13 domains, and 23 languages. Experiments with five frontier models show that no single model dominates across benchmark families and that cross-language disparities extend well beyond task success. Lower-resource languages exhibit distinct failure patterns, with larger shares of tool-use and control-flow errors rather than answer-quality errors alone, pointing to gaps in reliable task execution across the resource levels of these languages. On the same tasks, agents in low-resource languages also consume substantially more tokens than in English (up to roughly twice the input) without proportional increases in interaction length, and language consistency degrades further on tasks requiring structured output, where switches are directed overwhelmingly toward English. We believe BabelArena provides a foundation for advancing research on reliable and efficient multilingual agents.

Every Monday
Get next week’s papers.
Subscribe on Substack