On CelebA with high cheekbones as target and sex as sensitive attribute, zero-shot CLIP shows substantial bias measured by EOD. CLIP ResNet-50 achieves an EOD of 5.8% while CLIP ViT-L/14 achieves 2.8%, indicating that the model's predictions for cheekbones are systematically different for males versus females. This bias arises from a genuine correlation between the two attributes rather than a spurious one, and no existing debiasing baseline fully eliminates it.
Evidence
observational
Key metric
EOD: CLIP ResNet-50 5.8%, CLIP ViT-L/14 2.8% (CelebA, high cheekbones target, sex sensitive)