paper/arxiv/2609.03611
Robotics Dataset
FailBench is a benchmark for evaluating vision-language models on robot failure detection, comprising 2,197 manipulation attempts from 14 public sources, with 75% naturally occurring failures. It assesses 13 VLM-based detectors, revealing that general-purpose VLMs outperform fine-tuned ones and that performance degrades on contact-intensive tasks.
Catalog metadata
- Repository owner
- Zaruhi Navasardyan, Tatul Danielyan, Hrant Davtyan
- Source
- Research paper
- Task
- FailBench: How Reliable are VLMs at Judging Robot Task Success?
- Cameras
- 0
Browse the catalog to compare robots, tasks, sensors, and licenses. Sign in to open protected download and model links.