IC-456VLM decoders achieve near-random accuracy on VALSE image-sentence alignment while pairwise accuracy is much higher, indicating reliance on linguistic priors

Letitia Parcalabescu, Anette Frank

SourceDo Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?

The paper benchmarks four 7B VLM decoders on all VALSE instruments using both pairwise accuracy (accr) and overall accuracy (acc). In the acc setting, where the model must judge whether a single caption matches an image without seeing the foil, bakllava scores 50±0, llava-next-mistral 56±5, and llava-next-vicuna 59±9, all at or near the 50% random baseline. mPLUG-Owl3 reaches 76±10. In contrast, pairwise accuracy (accr), where both caption and foil are presented, is substantially higher: 83±8, 85±9, 79±10, and 88±6 respectively. The gap between accr and acc indicates the models exploit linguistic differences between caption and foil rather than genuinely understanding the image.

Evidence
correlational
Key metric
acc: bakllava 50±0, lv-mistral 56±5, lv-vicuna 59±9, mplug-owl3 76±10; accr: bakllava 83±8, lv-mistral 85±9, lv-vicuna 79±10, mplug-owl3 88±6
Caveat
Results are zero-shot; the paper notes that decoders of 2024 perform better than encoders of 2019-2021 in accr but not in acc. mPLUG-Owl3, the most recent model, does not exhibit the same linguistic bias pattern.
Model
BakLLaVA, LLaVA-NeXT / LLaVA 1.6 LLaVA-Next-Mistral, LLaVA-Next-Vicuna, mPLUG-Owl3
Concepts
Failure mode, Shortcut
Datasets
VALSE [eval], Foil It [eval]
Methods
MM-SHAP [supporting]
Related work
VALSE [builds-on]
Related findings
IC-457, IC-458
Extraction
automatic-extraction