paper/arxiv/2609.03611

Robotics Dataset

FailBench is a benchmark for evaluating vision-language models on robot failure detection, comprising 2,197 manipulation attempts from 14 public sources, with 75% naturally occurring failures. It assesses 13 VLM-based detectors, revealing that general-purpose VLMs outperform fine-tuned ones and that performance degrades on contact-intensive tasks.

Catalog metadata

Repository owner
Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan
Source
Research paper
Task
FailBench: How Reliable are VLMs at Judging Robot Task Success?
Cameras
0

Browse the catalog to compare robots, tasks, sensors, and licenses. Sign in to open protected download and model links.

Explore more robotics datasets and models