IC-1389Video-language models do not significantly outperform image-language models on temporal reasoning tasks in VILMA

Ilker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna, Emre Can Acikgoz, Letitia Parcalabescu, Iacer Calixto, Anette Frank, Albert Gatt, Aykut Erdem, Erkut Erdem

SourceViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

The paper evaluates 16 VidLMs, 2 ILMs (CLIP, BLIP-2), and 2 unimodal baselines (GPT-2, OPT) on five temporal reasoning tests (action counting, situation awareness, change of state, rare actions, spatial relations) using pairwise ranking accuracy. In the majority of tasks, VidLMs deliver performance levels closely resembling those of ILMs. In counting, situation awareness, and change of state, many VidLMs do not show notably higher performance than the random baseline (23.8 avg). The best ILM (BLIP-2, 57.5 avg P+T) matches or exceeds most VidLMs, with only Video-LLaMA (57.3) and InternVideo (54.2) approaching it.

Evidence
correlational
Key metric
Average P+T scores: BLIP-2 57.5, CLIP 52.9, Video-LLaMA 57.3, InternVideo 54.2, VindLU 49.5, FiT 48.2, Singularity 48.8, X-CLIP 47.8, MCQ 47.0, VioLET 46.3, Merlot Reserve 45.9, CLIP4Clip 52.1, VideoCLIP 38.9, UniVL 38.5, Otter 35.5, mPLUG-2 24.1, UniPerceiver 25.6, CLIPBERT 25.0, GPT-2 26.2, OPT 30.6, random 23.8
Caveat
The paper notes that foil generation introduces plausibility biases that may inflate unimodal and ILM scores in situation awareness and spatial relations. The evaluation uses a limited number of frames (k=4 or 8) for most models, which may further limit temporal reasoning.
Model
CLIPBERT, UniVL, VideoCLIP, CLIP4Clip, VioLET, X-CLIP, UniPerceiver, Merlot Reserve, VindLU, InternVideo, mPLUG-2, Otter, Video-LLaMA, CLIP / CLIP-ViT (LC), BLIP-2, GPT-2, OPT
Concepts
Failure mode
Datasets
QUVA [source], VidSitu [source], Something-Something V2 [source], YouCook2 [source], COIN [source], RareAct [source], STAR [source]
Related work
VALSE [context], SVO Probes [context], Park et al. 2022 (contrast sets for video-text) [context]
Related findings
IC-1390
Extraction
automatic-extraction