🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Evaluation

Agent Limitations Taxonomy

Free while signed in. Answers cite the passages they came from.

First page
Agent Limitations Taxonomy
The curator’s take

Benchmark scores keep climbing, yet the same agent failures resurface across otherwise unrelated evaluations, hidden behind the leaderboard. This University of Oxford work synthesizes 27 benchmark, taxonomy, and audit papers spanning 19 benchmarks into the first cross-cutting taxonomy of LLM-agent limitations.

Key points
01

Six failure clusters: The taxonomy names tool invocation and parameter errors, planning and constraint-satisfaction failures, long-horizon degradation from context accumulation, multi-agent coordination breakdowns, safety failures under adversarial or underspecified conditions, and measurement validity problems.

02

Failures compound nonlinearly: Reliability drops faster than task length grows, so strong sub-task scores do not add up to end-to-end success on longer tasks.

03

Scaffolding is not a fix: Adding scaffolding does not reliably improve reliability, undercutting the assumption that more orchestration automatically buys robustness.

04

Why it matters: By giving shared names to failures that leaderboards obscure, the taxonomy helps teams diagnose why agents break in production instead of trusting benchmark gains that do not transfer.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack