IC-464Video-LLaVA and Llama-VID systematically overestimate candidate VLM scores, assigning ratings near 4.00 across all visual dimensions and showing near-zero or negative agreement with a reference-guided agent-debate method

Ming Liu, Wensheng Zhang

SourceIs Your Video Language Model a Reliable Judge?

When used as judges to rate candidate VLM responses on the CVRR-ES benchmark, Video-LLaVA assigns scores close to 4.00 across nearly all 11 visual dimensions, and Llama-VID assigns scores typically above 3.70, regardless of the candidate model's actual quality. Their weighted Cohen's kappa agreement with the reference-guided LLM agent-debate method is near zero or negative (Video-LLaVA: -1.83 to 2.81; Llama-VID: -4.15 to 3.70), indicating they cannot discriminate between good and poor responses. This pattern is confirmed on a second dataset (VideoChatGPT), where Video-LLaVA's average kappa is only 1.72%.

Evidence
correlational
Key metric
Video-LLaVA scores 3.95-4.00 across dimensions; Llama-VID scores 3.76-4.03; Video-LLaVA kappa range -1.83 to 2.81; Llama-VID kappa range -4.15 to 3.70; VideoChatGPT dataset: Video-LLaVA average kappa 1.72%
Caveat
The reference (agent-debate) itself requires reference responses and is text-only (no visual input for GPT-3.5 agents), so the baseline may not be perfectly calibrated for video content.
Model
Video-LLaVA, Llama-VID
Concepts
Failure mode
Datasets
CVRR-ES [eval], VideoChatGPT [eval]
Methods
Weighted Cohen's Kappa [eval]
Related work
CVRR-ES [builds-on]
Related findings
IC-465, IC-466
Extraction
automatic-extraction