LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

Yujin Zhou, Mingxuan Zheng, Yike Guo, Sirui Han and colleagues at HKUST release LexAgentHallu, a 3,414-instance benchmark that annotates where along a legal agent's trajectory a hallucination originates, under a 7-category, 27-subclass taxonomy.
Ask this paper
Agentic hallucination, not answer hallucination: Tool-call and reasoning errors cascade into fabricated holdings and miscited authority. Existing legal benchmarks evaluate single-turn QA with outcome-level metrics, and agentic hallucination benchmarks lack legal-specific diagnostics.
Scale and construction: 3,414 instances across 17 legal categories and 6 task types, built through a four-stage expert-in-the-loop pipeline.
Dual-layer taxonomy: 7 high-level categories and 27 fine-grained subclasses covering substantive legal errors and agent-procedural failures separately, with metrics that localize where along the execution path each failure occurs.
Right-Answer-Wrong-Reason: Across 18 proprietary and open-source agents, the evaluation finds agents reaching correct answers through unsupported reasoning, which outcome-level scoring counts as success.
Failures cluster: Hallucination subclasses cluster rather than scatter, forming distinct profiles by agentic framework, by legal task, and by legal category, so a given framework has a characteristic failure shape.
Abstract
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and miscited authority. However, existing legal benchmarks evaluate only single-turn QA with outcome-level metrics, while agentic hallucination benchmarks lack legal-specific diagnostic capability. Neither answers to what extent and how a legal agent hallucinates along its trajectory. To address these limitations, we introduce LexAgentHallu, a legal agentic hallucination benchmark designed to evaluate to what extent and how legal agents fail along multi-step trajectories. Built through a four-stage expert-in-the-loop pipeline, LexAgentHallu contains 3414 instances across 17 legal categories and 6 task types. Each instance is annotated under a dual-layer hallucination taxonomy of 7 high-level categories and 27 fine-grained subclasses, covering both substantive errors and agent-procedural failures. We further design fine-grained metrics that quantify to what extent and localize how each failure occurs along an agent's execution path. Our evaluation across 18 proprietary and open-source agents uncovers a Right-Answer-Wrong-Reason effect and reveals that hallucination subclasses cluster rather than scatter, forming distinct agentic framework, legal task, and category profiles. These findings, invisible to outcome-level evaluation, validate the diagnostic power of LexAgentHallu for evaluating agentic hallucination in law.