In one qualitative CLIP explanation, FIXLIP assigns the strongest interaction to the caption token doll and an image patch containing the printed word dollar. The authors describe this as a case where one could say the model is right for the wrong reasons.
Evidence
observational
Caveat
This is one qualitative interaction map. The paper reports neither its frequency nor a targeted intervention removing the printed word, so it supports a local explanation of this prediction, not a general causal claim that printed text drives CLIP matching.