IC-040Information imbalance triggers both the modality gap and object bias in contrastive VLMs

Simon Schrodi, David T. Hoffmann, Max Argus, Volker Fischer, Thomas Brox

SourceTwo Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

The paper hypothesizes and validates that an information imbalance between image and text (where images contain more information than their captions) is the main cause for both the modality gap and object bias. Experiments on a synthetic dataset (MAD) and on real data (CC12M) show that reducing the information imbalance (by including more attributes in captions) decreases the modality gap and object bias while improving performance.

Evidence
interventional
Key metric
Modality gap (l2m & rmg) and object bias (MOAD) decrease as information imbalance is reduced; performance (accuracy and r@1) improves.
Caveat
Experiments were conducted on small- to medium-scale datasets (MAD, CC3M, CC12M) and may not fully represent the dynamics on larger datasets like LAION.
Model
CLIP / CLIP-ViT (LC)
Concepts
Failure mode
Datasets
CC12M [train], Conceptual 3M / CC3M [train]
Extraction
automatic-extraction