IC-835CLIP zero-shot predictions exhibit large worst-group accuracy gaps due to spurious correlations on Waterbirds and CelebA

Sepehr Dehdashtian, Lan Wang, Vishnu Boddeti

SourceFairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs

On Waterbirds (bird type vs background) and CelebA (blonde hair vs sex), zero-shot CLIP shows large gaps between average and worst-group accuracy. CLIP ViT-L/14 achieves only 45.3% worst-group accuracy on Waterbirds (gap 39.1%) and 72.8% on CelebA (gap 14.9%). CLIP ResNet-50 is worse on Waterbirds at 39.6% worst-group (gap 37.7%). The model relies on the background or sex as a proxy for the target attribute, a correlation that does not hold causally.

Evidence
observational
Key metric
Zero-shot WG/avg/gap: CLIP ViT-L/14 Waterbirds 45.3/84.4/39.1, CelebA 72.8/87.6/14.9; CLIP ResNet-50 Waterbirds 39.6/77.3/37.7, CelebA 75.9/82.3/6.4
Model
CLIP / CLIP-ViT (LC)
Concepts
Shortcut
Datasets
Waterbirds [eval], CelebA [eval]
Related work
Contrastive adapter [compared-to], Orth-Cali [compared-to], ERM adapter [compared-to]
Related findings
IC-834, IC-836, IC-837
Extraction
automatic-extraction