IC-518GPT-4o achieves 62.54% overall accuracy on MMWorld, the best among 15 MLLMs, while four open-source models perform below the 26.31% random-choice baseline
Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, Xin Eric Wang
The paper evaluates 15 MLLMs on the MMWorld benchmark, which spans 7 disciplines and 69 subdisciplines with multi-faceted reasoning questions (explanation, counterfactual thinking, future prediction, domain expertise, temporal understanding, attribution, procedure). GPT-4o leads with 62.54% average accuracy, followed by Claude-3.5-Sonnet at 54.54%. Four open-source models (Video-LLaMA-2-13B at 14.03%, LWM-1M-JAX at 15.39%, Otter-7B at 14.99%, X-Instruct-Blip-7B at 21.36%) score below the random-choice baseline of 26.31%. The open-source Video-LLaVA-7B outperforms GPT-4V and Gemini Pro on embodied tasks (63.17% vs 55.48% and 43.59%) and leads on temporal understanding (34.45% vs 27.17% and 24.65%).
Evidence
correlational
Key metric
GPT-4o 62.54% ±0.79 overall; random choice 26.31%; Video-LLaVA-7B embodied tasks 63.17% ±1.44 vs GPT-4V 55.48% ±2.70 and Gemini Pro 43.59% ±0.33; Video-LLaVA temporal understanding 34.45% ±1.19 vs GPT-4V 27.17% ±1.00