Agent Limitations Taxonomy

Benchmark scores keep climbing, yet the same agent failures resurface across otherwise unrelated evaluations, hidden behind the leaderboard. This University of Oxford work synthesizes 27 benchmark, taxonomy, and audit papers spanning 19 benchmarks into the first cross-cutting taxonomy of LLM-agent limitations.
Ask this paper
Six failure clusters: The taxonomy names tool invocation and parameter errors, planning and constraint-satisfaction failures, long-horizon degradation from context accumulation, multi-agent coordination breakdowns, safety failures under adversarial or underspecified conditions, and measurement validity problems.
Failures compound nonlinearly: Reliability drops faster than task length grows, so strong sub-task scores do not add up to end-to-end success on longer tasks.
Scaffolding is not a fix: Adding scaffolding does not reliably improve reliability, undercutting the assumption that more orchestration automatically buys robustness.
Why it matters: By giving shared names to failures that leaderboards obscure, the taxonomy helps teams diagnose why agents break in production instead of trusting benchmark gains that do not transfer.