IC-465GPT-4o achieves the highest weighted Cohen's kappa agreement with a reference-guided agent-debate method among all VLM judges, with scores exceeding 50 in several visual dimensions
Among the five VLMs tested as judges (Video-LLaVA, Llama-VID, GPT-4o mini, InternVL2, GPT-4o), GPT-4o shows the strongest agreement with the reference-guided LLM agent-debate method. Its average kappa reaches 60.38 for social context and 55.18 for visual context, substantially above all other VLMs. The paper notes a consistent trend: the better a judge performs on the CVRR-ES benchmark itself, the higher its kappa as a judge, suggesting that content understanding is a prerequisite for reliable evaluation.
The agent-debate reference is text-only (GPT-3.5 and GPT-4o in text mode with reference answers), so agreement with it may not fully capture visual evaluation quality. The paper acknowledges reliance on a specific set of models and datasets.