Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
AHA dataset
anchor
Note
introduced by the paper that uses it, so the anchor is that paper; the dataset has no separate release of its own that the source prints
Findings
IC-005
GPT-4o underperforms AHA and other VLMs in detecting and reasoning about robotic manipulation failures across multiple datasets.
[eval]