IC-519MLLMs exhibit different skill sets than humans, correctly answering expert-level questions that all three human annotators miss while failing on easy questions humans answer correctly
Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, Xin Eric Wang
The paper defines four difficulty levels based on 3 Amazon Turkers' performance per question (easy: 3/3 correct, medium: 2/3, hard: 1/3, expert: 0/3). While MLLMs generally show decreasing accuracy with increasing difficulty, they also correctly answer expert-level questions that all three humans got wrong, particularly in business and health & medicine. Conversely, they sometimes fail on easy questions that all humans answered correctly. GPT-4V scores 92% on expert-level business questions but only 60% on easy-level art & sports questions, and 87% on expert-level health & medicine but 46% on easy-level embodied tasks.
Evidence
correlational
Key metric
GPT-4V: 92% expert-level business, 60% easy-level art & sports; 87% expert-level health & medicine, 46% easy-level embodied tasks
Caveat
Difficulty levels are defined by only 3 Turkers per question, which may not capture the full range of human difficulty.