The paper evaluates free-text explanations from VLMs for correct predictions in the triplet setup. Human annotators judged whether the explanation logically justifies why the correct hypothesis is more plausible. Humans achieve 95% valid explanations, while the best VLM (Claude-Sonnet-3.5) only achieves 50%. GPT-4o achieves 44% and 23% in separate and combined image strategies, respectively. Many models score below 40%, indicating that even correct predictions are often based on shallow or incorrect reasoning.
Evidence
correlational
Key metric
Human explanation validity 95%; Claude-Sonnet-3.5 separate images 50%, GPT-4o separate images 23%, GPT-4-Vision 39%, Gemini-1.5-Pro 38%
Caveat
Explanation evaluation is subjective and uses a binary validity judgment; the reference explanations may not cover all possible valid reasoning paths.