IC-836CLIP zero-shot predictions exhibit demographic bias on FairFace when using attribute-unrelated text prompts

Sepehr Dehdashtian, Lan Wang, Vishnu Boddeti

SourceFairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs

On the FairFace dataset, using ten text prompts unrelated to facial or sensitive attributes (e.g., 'a photo of a criminal person'), zero-shot CLIP shows high maxSkew@1000 for both sex and race. CLIP ViT-B/32 achieves maxSkew of 0.206 (sex) and 0.743 (race), while CLIP ViT-L/14 achieves 0.206 (sex) and 0.768 (race). This indicates that CLIP's representations encode race and sex information that leaks into predictions for unrelated target attributes.

Evidence
observational
Key metric
maxSkew@1000: CLIP ViT-B/32 sex 0.206, race 0.743; CLIP ViT-L/14 sex 0.206, race 0.768
Model
CLIP / CLIP-ViT (LC)
Concepts
Shortcut
Datasets
FairFace [eval]
Related work
Orth-Proj [compared-to], Orth-Cali [compared-to]
Related findings
IC-834, IC-835, IC-837
Extraction
automatic-extraction