IC-519MLLMs exhibit different skill sets than humans, correctly answering expert-level questions that all three human annotators miss while failing on easy questions humans answer correctly

Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, Xin Eric Wang

SourceMMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

The paper defines four difficulty levels based on 3 Amazon Turkers' performance per question (easy: 3/3 correct, medium: 2/3, hard: 1/3, expert: 0/3). While MLLMs generally show decreasing accuracy with increasing difficulty, they also correctly answer expert-level questions that all three humans got wrong, particularly in business and health & medicine. Conversely, they sometimes fail on easy questions that all humans answered correctly. GPT-4V scores 92% on expert-level business questions but only 60% on easy-level art & sports questions, and 87% on expert-level health & medicine but 46% on easy-level embodied tasks.

Evidence
correlational
Key metric
GPT-4V: 92% expert-level business, 60% easy-level art & sports; 87% expert-level health & medicine, 46% easy-level embodied tasks
Caveat
Difficulty levels are defined by only 3 Turkers per question, which may not capture the full range of human difficulty.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, GPT-4o, Claude 3.5 Sonnet, Gemini Pro, Video-LLaVA Video-LLaVA-7B, Video-Chat-7B, Video-ChatGPT Video-ChatGPT-7B, ImageBind-LLM-7B, PandaGPT-7B, Chat-UniVi-7B, Video-LLaMA-2-13B, X-InstructBLIP-7B, LWM-1M-JAX, Otter-7B, mPLUG-Owl mPLUG-Owl-7B
Related findings
IC-518, IC-520, IC-521
Extraction
automatic-extraction