Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1389
Video-language models do not significantly outperform image-language models on temporal reasoning tasks in VILMA
IC-1390
Proficiency tests reveal that a substantial portion of correct main-test predictions by VidLMs and ILMs are spurious rather than reflecting robust understanding