IC-039Object bias in CLIP and SigLIP is not correlated with performance on attribute tasks

Simon Schrodi, David T. Hoffmann, Max Argus, Volker Fischer, Thomas Brox

SourceTwo Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

The paper measures object bias in contrastive VLMs using a new metric, Moad. They find that models pre-trained on large-scale data tend to have a lower object bias than those on medium-scale data. However, there is no clear correlation between object bias and performance on attribute recognition tasks. The paper finds that performance improvements on object tasks correlate with improvements on attribute tasks, suggesting that general model improvement also benefits attribute recognition.

Evidence
correlational
Key metric
No clear correlation found; medium-to-strong correlations between object and attribute performance (Kendall’s τ: 39.0 to 78.1).
Caveat
The metric Moad is a new measure and its relationship to other forms of bias is not fully explored.
Model
CLIP / CLIP-ViT (LC), SigLIP
Datasets
MIT-States [eval], UT-Zappos [eval], MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [eval], ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval]
Extraction
automatic-extraction