Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
E.T.Bench
anchor
Findings
IC-049
GPT-4o shows strong performance on some E.T.Bench event-level tasks (RVQ: 57.7, VHD: 56.9) but very weak performance on others (EPM: 4.5, TAL: 20.0)
[eval]
IC-051
Qwen2-VL (7B) achieves 0.0 on DVC, DVC SLC, and TEM tasks on E.T.Bench
[eval]