IC-1356GPT-3.5 turbo achieves over 90% agreement with human judgments when evaluating VLM responses on open-set questions

Yuanfeng Ji, Chongjian GE, Weikai Kong, Enze Xie, Zhengying Liu, Zhenguo Li, Ping Luo

SourceLarge Language Models as Automated Aligners for benchmarking Vision-Language Models

To validate the LLM-as-judge approach, the authors had both human judges and GPT-3.5 turbo independently assess 100 randomly selected open-set questions per capacity (perception, planning, value) from the eight VLMs. GPT-3.5 turbo's correctness judgments align with human judgments at an average accuracy of over 90%, as shown in a box plot across the three capacities. This validates the use of GPT-3.5 as an automated evaluator for the benchmark.

Evidence
correlational
Key metric
Average accuracy of over 90% in providing correct judgments (Figure 5, box plot across perception, planning, and value capacities)
Caveat
Validation is limited to 100 questions per capacity (300 total) and only covers open-set questions; closed-set (multiple-choice) judgment accuracy is not separately validated against humans.
Model
GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo
Datasets
Auto-Bench [eval]
Related findings
IC-1355
Extraction
automatic-extraction