Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
DynaMath
Findings
IC-276
All 14 evaluated VLMs show a large gap between average-case and worst-case accuracy on DynaMath variants, with worst-case at or below 50% of average-case, and the failures are systematic rather than random
[eval]
IC-277
Open-source VLMs show a clear scaling trend in both average accuracy and reasoning robustness on DynaMath, with larger models performing substantially better
[eval]
IC-278
Claude-3.5 Sonnet and GPT-4o exhibit a memorization failure mode, outputting the same answer regardless of visual parameter changes in the problem
[eval]