IC-520Temporal reasoning performance drops significantly across all 15 MLLMs when video frames are shuffled or reduced to one-fifth of the original count

Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang, Lijuan Wang, Xin Eric Wang

SourceMMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

The paper tests 15 MLLMs on temporal understanding questions under three conditions: original videos, shuffled video frames, and videos reduced to 1/5 of their original frame count. All models show performance degradation under both perturbations. GPT-4o drops from 40.90% to 35.11% (shuffled) and 32.19% (reduced). Proprietary models (GPT-4o, GPT-4V) show better resilience than most open-source models. The largest drops are seen in Video-LLaVA (34.45% to 18.47% shuffled) and Otter-7B (9.52% to 3.25% shuffled), indicating high sensitivity to temporal coherence.

Evidence
correlational
Key metric
GPT-4o: 40.90% original, 35.11% shuffled, 32.19% reduced; Video-LLaVA: 34.45% original, 18.47% shuffled; Otter-7B: 9.52% original, 3.25% shuffled
Model
GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, Claude 3.5 Sonnet, Gemini Pro, Video-LLaVA Video-LLaVA-7B, Video-Chat-7B, Video-ChatGPT Video-ChatGPT-7B, ImageBind-LLM-7B, PandaGPT-7B, Chat-UniVi-7B, Video-LLaMA-2-13B, X-InstructBLIP-7B, LWM-1M-JAX, Otter-7B, mPLUG-Owl mPLUG-Owl-7B
Concepts
Failure mode
Related findings
IC-518, IC-519, IC-521
Extraction
automatic-extraction