Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Measuring Vision-Language STEM Skills of Neural Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1489
State-of-the-art foundation models (CLIP, GPT-3.5-turbo, and others) score well below elementary students on multimodal K-12 STEM questions
IC-1490
Zero-shot CLIP is overconfident on STEM questions, with softmax confidence loosely related to actual accuracy
IC-1491
CLIP zero-shot performance on STEM saturates across model sizes, with only 3.6 points of variation from smallest to largest variant