Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Large Language Models as Automated Aligners for benchmarking Vision-Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1355
Instruction-tuned VLMs fail to follow multiple-choice format in reasoning questions, with InstructBLIP frequently returning blank responses
IC-1356
GPT-3.5 turbo achieves over 90% agreement with human judgments when evaluating VLM responses on open-set questions