FX-001Pairwise Banzhaf interactions explain CLIP similarity more faithfully than single-score methods

Hubert Baniecki, Maximilian Muschalik, Fabian Fumagalli, Barbara Hammer, Eyke Hüllermeier, Przemyslaw Biecek

SourceFIXLIP: Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions

Most explanation methods for CLIP hand you one importance score per image patch and per word. FIXLIP also scores every pair, so you can see which word interacts with which part of the picture. These pairwise explanations rebuild CLIP's similarity score faithfully and beat the single-score methods GAME and Grad-ECLIP by a wide margin once more than one object is in the picture.

Evidence
correlational
Key metric
pointing game 0.83 / 0.81 / 0.83 / 0.85 for 1-4 objects (weighted Banzhaf, p = 0.7), versus GAME 0.61 / 0.43 / 0.33 / 0.28
Caveat
A rival pairwise method, exCLIP, scores higher still on the multi-object cases.
Model
CLIP / CLIP-ViT (LC) CLIP ViT-B/32
Concepts
Feature interaction, Explanation faithfulness
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [eval], ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval]
Methods
Weighted Banzhaf interaction index [primary], Faith-Shap [compared-to], Pointing game [eval], Area between insertion and deletion curves [eval]
Related work
GAME (Generic Attention-model Explainability) [compared-to], Grad-ECLIP [compared-to], exCLIP (second-order attributions for caption-image interactions) [compared-to]
Related findings
FX-002, FX-003
Extraction
manual-extraction