AgentBoard
Free while signed in. Answers cite the passages they came from.

AgentBoard is a benchmark and open-source evaluation framework for analytically evaluating LLM agents beyond the usual pass/fail metrics.
Fine-grained progress metric: Measures progress across sub-goals rather than a single end-of-trajectory success flag, giving much richer signal about where agents fail.
Diverse task suite: Covers 9 task categories spanning embodied AI, web, tool use, and game environments, designed to stress different agent capabilities.
Interactive visualization: Ships with a GUI for visualizing trajectories, enabling qualitative analysis of failure modes rather than relying on aggregate scores.
Robust agent insights: Reveals systematic weaknesses in current LLM agents - long-horizon planning, memory, and dealing with partially observable environments - that aggregate metrics tend to hide.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack