IC-050GPT-4o achieves 53.33 overall on Event-Bench with strong event description (57.50) and counter reasoning (63.44) but weaker episodic reasoning (37.33)
The paper evaluates GPT-4o on Event-Bench, a benchmark for causal reasoning over videos. GPT-4o achieves the highest overall score among all models tested (53.33), with particularly strong performance on event description (57.50 overall) and counterfactual reasoning (63.44). However, its episodic reasoning score (37.33) is notably lower than its other causal reasoning sub-scores, indicating a relative weakness in that specific reasoning type.