IC-002Even when VLMs select the correct hypothesis, their explanations are often invalid or unhelpful

Mor Ventura, Michael Toker, Nitay Calderon, Zorik Gekhman, Yonatan Bitton, Roi Reichart

SourceNL-Eye: Abductive NLI For Images

The paper evaluates free-text explanations from VLMs for correct predictions in the triplet setup. Human annotators judged whether the explanation logically justifies why the correct hypothesis is more plausible. Humans achieve 95% valid explanations, while the best VLM (Claude-Sonnet-3.5) only achieves 50%. GPT-4o achieves 44% and 23% in separate and combined image strategies, respectively. Many models score below 40%, indicating that even correct predictions are often based on shallow or incorrect reasoning.

Evidence
correlational
Key metric
Human explanation validity 95%; Claude-Sonnet-3.5 separate images 50%, GPT-4o separate images 23%, GPT-4-Vision 39%, Gemini-1.5-Pro 38%
Caveat
Explanation evaluation is subjective and uses a binary validity judgment; the reference explanations may not cover all possible valid reasoning paths.
Model
Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, GPT-4o, Claude 3.5 Sonnet, Claude 3 Opus, LLaVA-NeXT / LLaVA 1.6, Fuyu
Concepts
Failure mode, Explanation faithfulness
Datasets
NL-Eye [eval]
Extraction
automatic-extraction