Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-518
GPT-4o achieves 62.54% overall accuracy on MMWorld, the best among 15 MLLMs, while four open-source models perform below the 26.31% random-choice baseline
IC-519
MLLMs exhibit different skill sets than humans, correctly answering expert-level questions that all three human annotators miss while failing on easy questions humans answer correctly
IC-520
Temporal reasoning performance drops significantly across all 15 MLLMs when video frames are shuffled or reduced to one-fifth of the original count
IC-521
MLLMs show asymmetric modality-specific perception, with Gemini Pro achieving 69.97% on visual-only questions but only 24.45% on audio-only, while Video-Chat outperforms ChatUniVi on audio despite worse visual scores