IC-518GPT-4o achieves 62.54% overall accuracy on MMWorld, the best among 15 MLLMs, while four open-source models perform below the 26.31% random-choice baseline

Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, Xin Eric Wang

SourceMMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

The paper evaluates 15 MLLMs on the MMWorld benchmark, which spans 7 disciplines and 69 subdisciplines with multi-faceted reasoning questions (explanation, counterfactual thinking, future prediction, domain expertise, temporal understanding, attribution, procedure). GPT-4o leads with 62.54% average accuracy, followed by Claude-3.5-Sonnet at 54.54%. Four open-source models (Video-LLaMA-2-13B at 14.03%, LWM-1M-JAX at 15.39%, Otter-7B at 14.99%, X-Instruct-Blip-7B at 21.36%) score below the random-choice baseline of 26.31%. The open-source Video-LLaVA-7B outperforms GPT-4V and Gemini Pro on embodied tasks (63.17% vs 55.48% and 43.59%) and leads on temporal understanding (34.45% vs 27.17% and 24.65%).

Evidence
correlational
Key metric
GPT-4o 62.54% ±0.79 overall; random choice 26.31%; Video-LLaVA-7B embodied tasks 63.17% ±1.44 vs GPT-4V 55.48% ±2.70 and Gemini Pro 43.59% ±0.33; Video-LLaVA temporal understanding 34.45% ±1.19 vs GPT-4V 27.17% ±1.00
Model
GPT-4o, Claude 3.5 Sonnet, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, Gemini Pro, Video-LLaVA Video-LLaVA-7B, Video-Chat-7B, Chat-UniVi-7B, mPLUG-Owl mPLUG-Owl-7B, Video-ChatGPT Video-ChatGPT-7B, PandaGPT-7B, ImageBind-LLM-7B, X-InstructBLIP-7B, LWM-1M-JAX, Otter-7B, Video-LLaMA-2-13B
Methods
GPT-4-32k [eval]
Related work
MVBench [context], TempCompass [context], Perception Test [context]
Related findings
IC-519, IC-520, IC-521
Extraction
automatic-extraction