IC-049GPT-4o shows strong performance on some E.T.Bench event-level tasks (RVQ: 57.7, VHD: 56.9) but very weak performance on others (EPM: 4.5, TAL: 20.0)

Yongxin Guo, Jingyu Liu, Mingda Li, Qingbin Liu, Xi Chen, Xiaoying Tang

SourceTRACE: Temporal Grounding Video LLM via Causal Event Modeling

The paper evaluates GPT-4o on the E.T.Bench benchmark, which covers multiple event-level video understanding tasks. GPT-4o achieves high scores on relational video question answering (RVQ: 57.7) and video highlight detection (VHD: 56.9), but scores near zero on event prediction matching (EPM: 4.5) and temporal action localization (TAL: 20.0). This pattern suggests GPT-4o's video understanding is uneven across event-level task types.

Evidence
correlational
Key metric
E.T.Bench: RAR 27.8, ECA 27.3, RVQ 57.7, TVG 40.4, EPM 4.5, TAL 20.0, EVS 17.6, VHD 56.9, DVC 46.9, DVC SLC 22.3, TEM 23.1, GVQ 14.9
Model
GPT-4o
Concepts
Failure mode
Datasets
E.T.Bench [eval]
Related findings
IC-048, IC-050, IC-051
Extraction
automatic-extraction