Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Auto-Bench
Findings
IC-1355
Instruction-tuned VLMs fail to follow multiple-choice format in reasoning questions, with InstructBLIP frequently returning blank responses
[eval]
IC-1356
GPT-3.5 turbo achieves over 90% agreement with human judgments when evaluating VLM responses on open-set questions
[eval]