To validate the LLM-as-judge approach, the authors had both human judges and GPT-3.5 turbo independently assess 100 randomly selected open-set questions per capacity (perception, planning, value) from the eight VLMs. GPT-3.5 turbo's correctness judgments align with human judgments at an average accuracy of over 90%, as shown in a box plot across the three capacities. This validates the use of GPT-3.5 as an automated evaluator for the benchmark.
Evidence
correlational
Key metric
Average accuracy of over 90% in providing correct judgments (Figure 5, box plot across perception, planning, and value capacities)
Caveat
Validation is limited to 100 questions per capacity (300 total) and only covers open-set questions; closed-set (multiple-choice) judgment accuracy is not separately validated against humans.