IC-038Few embedding dimensions drive the modality gap in CLIP and SigLIP

Simon Schrodi, David T. Hoffmann, Max Argus, Volker Fischer, Thomas Brox

SourceTwo Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models

The paper analyzes the embedding space of off-the-shelf CLIP and SigLIP models. It measures the mean difference per embedding dimension between the image and text modalities. Most dimensions show similar means, but a small subset shows large differences. Two of these dimensions are sufficient to perfectly separate the image and text embeddings.

Evidence
observational
Key metric
Two dimensions suffice to perfectly separate the modalities.
Caveat
The finding is based on analysis of the embedding geometry, not on downstream task performance.
Model
CLIP / CLIP-ViT (LC), SigLIP
Concepts
Linear representation
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [eval], ImageNet-1k / ImageNet / ImageNet-1k-val / ImageNet-Val [eval], CIFAR-10 [eval], CIFAR-100 [eval]
Methods
Normalized Kendall-tau distance / Normalized Kendall-τ distance [supporting]
Extraction
automatic-extraction