Agent Limitations Taxonomy
Free while signed in. Answers cite the passages they came from.

Benchmark scores keep climbing, yet the same agent failures resurface across otherwise unrelated evaluations, hidden behind the leaderboard. This University of Oxford work synthesizes 27 benchmark, taxonomy, and audit papers spanning 19 benchmarks into the first cross-cutting taxonomy of LLM-agent limitations.
Six failure clusters: The taxonomy names tool invocation and parameter errors, planning and constraint-satisfaction failures, long-horizon degradation from context accumulation, multi-agent coordination breakdowns, safety failures under adversarial or underspecified conditions, and measurement validity problems.
Failures compound nonlinearly: Reliability drops faster than task length grows, so strong sub-task scores do not add up to end-to-end success on longer tasks.
Scaffolding is not a fix: Adding scaffolding does not reliably improve reliability, undercutting the assumption that more orchestration automatically buys robustness.
Why it matters: By giving shared names to failures that leaderboards obscure, the taxonomy helps teams diagnose why agents break in production instead of trusting benchmark gains that do not transfer.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack