Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1288
GPT-4, ChatGPT, and GPT-4V fail to close the human-machine gap on Bongard-OpenWorld, with InstructBLIP captions differentially degrading ChatGPT while improving GPT-4
IC-1289
OpenFlamingo and Otter achieve near-chance accuracy on Bongard-OpenWorld, indicating inability to perform multi-image reasoning
IC-1290
CLIP, DINO, and DINOv2 as zero-shot natural baselines score below the 50% chance level on Bongard-OpenWorld due to adversarial query selection