Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
NL-Eye
anchor
Note
introduced by the paper that uses it, so the anchor is that paper; the dataset has no separate release of its own that the source prints
Findings
IC-001
Vision-language models perform near chance on the NL-Eye visual abductive reasoning benchmark
[eval]
IC-002
Even when VLMs select the correct hypothesis, their explanations are often invalid or unhelpful
[eval]