FX-002FIXLIP gives SigLIP-2 higher pointing-game recognition than CLIP at ViT-B/32 and ViT-B/16

Hubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, Eyke Hüllermeier, Przemyslaw Biecek

SourceFIXLIP: Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions

In Table 2, FIXLIP explanations of SigLIP-2 obtain higher pointing-game recognition than CLIP at both reported patch configurations, ViT-B/32 and ViT-B/16. The authors interpret this as SigLIP-2 being more faithfully explainable through cross-modal interactions and suggest that smaller models may be more faithfully explainable than larger ones.

Evidence
correlational
Key metric
pointing game recognition for 1 to 4 objects, ViT-B/32: SigLIP-2 0.90 / 0.90 / 0.89 / 0.89 against CLIP 0.83 / 0.81 / 0.82 / 0.85; ViT-B/16: SigLIP-2 0.86 / 0.88 / 0.87 / 0.87 against CLIP 0.81 / 0.81 / 0.81 / 0.82
Caveat
The comparison does not use exactly the same pointing-game cases: subgames containing goldfish or ipod are omitted for SigLIP models because their tokenizer splits those labels into multiple tokens. The authors offer no mechanism for the SigLIP-2 result. Their scale interpretation is not uniform across families: SigLIP scores no lower at ViT-L/16 than at ViT-B/16.
Model
CLIP / CLIP-ViT (LC) CLIP ViT-B/32, CLIP ViT-B/16, SigLIP ViT-B/16, ViT-L/16, SigLIP-2 ViT-B/32, ViT-B/16, ViT-L/16
Concepts
Explanation faithfulness
Datasets
ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval]
Methods
Weighted Banzhaf interaction index [primary], Pointing game [eval]
Related findings
FX-001
Extraction
automatic-extraction