Survey on Evaluation of LLM-based Agents
Free while signed in. Answers cite the passages they came from.

This work presents the first comprehensive overview of how to evaluate LLM-based agents, which differ significantly from traditional LLMs by maintaining memory, planning over multiple steps, using tools, and interacting with dynamic environments. The authors categorize and analyze the evaluation landscape across four axes: core agent capabilities, application-specific agent benchmarks, generalist agent evaluation, and supporting evaluation frameworks.
The paper organizes fundamental capabilities into four categories: (1) planning and multi-step reasoning (e.g., GSM8K, PlanBench, MINT), (2) function calling and tool use (e.g., ToolBench, BFCL, ToolSandbox), (3) self-reflection (e.g., LLF-Bench, LLM-Evolve), and (4) memory (e.g., MemGPT, ReadAgent, A-MEM, StreamBench). For each, it surveys benchmarks that test these competencies under increasingly realistic and complex scenarios.
In domain-specific evaluation, the authors highlight the rise of specialized agents in web navigation (e.g., WebShop, WebArena), software engineering (e.g., SWE-bench and its variants), scientific research (e.g., ScienceQA, AAAR-1.0, DiscoveryWorld), and dialogue (e.g., ABCD, τ-Bench, IntellAgent). These benchmarks typically include goal-oriented tasks, tool use, and policy adherence, often with simulations or human-in-the-loop data.
Generalist agent benchmarks like GAIA, AgentBench, OSWorld, and TheAgentCompany aim to evaluate flexibility across heterogeneous tasks, integrating planning, tool use, and real-world digital workflows. Benchmarks such as HAL aim to unify evaluation across multiple axes, including coding, reasoning, and safety.
The paper also covers evaluation frameworks like LangSmith, Langfuse, Vertex AI, and Galileo Agentic Evaluation, which allow stepwise, trajectory-based, and human-in-the-loop assessments. These systems enable real-time monitoring and debugging of agent behavior during development, often via synthetic data and LLM-as-a-judge pipelines.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack