Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
CVRR-ES
anchor
Findings
IC-464
Video-LLaVA and Llama-VID systematically overestimate candidate VLM scores, assigning ratings near 4.00 across all visual dimensions and showing near-zero or negative agreement with a reference-guided agent-debate method
[builds-on]
IC-464
Video-LLaVA and Llama-VID systematically overestimate candidate VLM scores, assigning ratings near 4.00 across all visual dimensions and showing near-zero or negative agreement with a reference-guided agent-debate method
[eval]
IC-465
GPT-4o achieves the highest weighted Cohen's kappa agreement with a reference-guided agent-debate method among all VLM judges, with scores exceeding 50 in several visual dimensions
[builds-on]
IC-465
GPT-4o achieves the highest weighted Cohen's kappa agreement with a reference-guided agent-debate method among all VLM judges, with scores exceeding 50 in several visual dimensions
[eval]
IC-466
GPT-4o's evaluation reliability degrades when used as the final judge in a collective thought pipeline that aggregates reviews from less reliable VLMs
[builds-on]
IC-466
GPT-4o's evaluation reliability degrades when used as the final judge in a collective thought pipeline that aggregates reviews from less reliable VLMs
[eval]