IC-520Temporal reasoning performance drops significantly across all 15 MLLMs when video frames are shuffled or reduced to one-fifth of the original count
Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, Xin Eric Wang
The paper tests 15 MLLMs on temporal understanding questions under three conditions: original videos, shuffled video frames, and videos reduced to 1/5 of their original frame count. All models show performance degradation under both perturbations. GPT-4o drops from 40.90% to 35.11% (shuffled) and 32.19% (reduced). Proprietary models (GPT-4o, GPT-4V) show better resilience than most open-source models. The largest drops are seen in Video-LLaVA (34.45% to 18.47% shuffled) and Otter-7B (9.52% to 3.25% shuffled), indicating high sensitivity to temporal coherence.