SourceTRACE: Temporal Grounding Video LLM via Causal Event Modeling
The paper evaluates GPT-4o on the E.T.Bench benchmark, which covers multiple event-level video understanding tasks. GPT-4o achieves high scores on relational video question answering (RVQ: 57.7) and video highlight detection (VHD: 56.9), but scores near zero on event prediction matching (EPM: 4.5) and temporal action localization (TAL: 20.0). This pattern suggests GPT-4o's video understanding is uneven across event-level task types.