Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
VideoChatGPT
Findings
IC-464
Video-LLaVA and Llama-VID systematically overestimate candidate VLM scores, assigning ratings near 4.00 across all visual dimensions and showing near-zero or negative agreement with a reference-guided agent-debate method
[eval]
IC-465
GPT-4o achieves the highest weighted Cohen's kappa agreement with a reference-guided agent-debate method among all VLM judges, with scores exceeding 50 in several visual dimensions
[eval]