🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 5 – Sep 5, 2026
Robotics · Training

FailBench: How Reliable are VLMs at Judging Robot Task Success?

First page
FailBench: How Reliable are VLMs at Judging Robot Task Success?
The curator’s take

Zaruhi Navasardyan, Tatul Danielyan and Hrant Davtyan at Metric AI Lab assemble 2,197 real manipulation attempts from 14 sources and find that vision-language models used as robot success detectors reach only 0.77 mean balanced accuracy, with fine-tuned detectors doing worse than general-purpose models.

Ask this paper

Key points
01

The benchmark avoids the usual staging problem: 75 percent of failures occur naturally and six of the real-world sources were not collected for failure detection at all, so the distribution is not curated toward detectable failures.

02

Fine-tuning hurts: models fine-tuned for failure detection consistently underperform general-purpose VLMs and even their own pretrained baselines, which is a direct argument against the standard recipe.

03

Performance splits by required visual evidence: near-saturation when the outcome depends on observable object motion, but below 0.60 balanced accuracy on contact-intensive assembly, alongside a systematic bias toward predicting success under ambiguous evidence that persists with more reasoning effort.

04

One intervention that works: spatially localizing and cropping the outcome-relevant region improves the top detector by 2.4 points with no training.

05

Why it matters for agents: these judgments are used as RL rewards, training-data filters, policy-ranking signals and retry triggers, so a 0.77 detector silently corrupts all four.

Abstract

Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated). In FailBench, 75% of failures occur naturally, and six real-world sources come from non-failure-detection datasets. Evaluating 13 VLM-based detectors, we find the best model achieves only 0.77 mean balanced accuracy. Notably, models fine-tuned for failure detection consistently underperform general-purpose VLMs and their own pretrained baselines. Performance depends heavily on required visual evidence: models approach saturation when outcomes depend on observable object motion, but degrade to near-chance (<0.60 balanced accuracy) on contact-intensive assembly tasks. Error analysis reveals a systematic bias toward predicting success under ambiguous evidence, which persists even with increased reasoning effort. Finally, we show that input-level intervention--spatially localizing and cropping outcome-relevant regions--improves the top detector by 2.4 percentage points without extra training.

Every Monday
Get next week’s papers.
Subscribe on Substack