FX-003FIXLIP's strongest interaction in one CLIP example links doll to an image patch reading dollar

Hubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, Eyke Hüllermeier, Przemyslaw Biecek

SourceFIXLIP: Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions

In one qualitative CLIP explanation, FIXLIP assigns the strongest interaction to the caption token doll and an image patch containing the printed word dollar. The authors describe this as a case where one could say the model is right for the wrong reasons.

Evidence
observational
Caveat
This is one qualitative interaction map. The paper reports neither its frequency nor a targeted intervention removing the printed word, so it supports a local explanation of this prediction, not a general causal claim that printed text drives CLIP matching.
Model
CLIP / CLIP-ViT (LC)
Concepts
Shortcut, Feature interaction
Methods
Weighted Banzhaf interaction index [primary]
Related work
Eyes Wide Shut? Exploring the visual shortcomings of multimodal LLMs [context]
Related findings
FX-001
Extraction
automatic-extraction