On Waterbirds (bird type vs background) and CelebA (blonde hair vs sex), zero-shot CLIP shows large gaps between average and worst-group accuracy. CLIP ViT-L/14 achieves only 45.3% worst-group accuracy on Waterbirds (gap 39.1%) and 72.8% on CelebA (gap 14.9%). CLIP ResNet-50 is worse on Waterbirds at 39.6% worst-group (gap 37.7%). The model relies on the background or sex as a proxy for the target attribute, a correlation that does not hold causally.