The paper analyzes the embedding space of off-the-shelf CLIP and SigLIP models. It measures the mean difference per embedding dimension between the image and text modalities. Most dimensions show similar means, but a small subset shows large differences. Two of these dimensions are sufficient to perfectly separate the image and text embeddings.
Evidence
observational
Key metric
Two dimensions suffice to perfectly separate the modalities.
Caveat
The finding is based on analysis of the embedding geometry, not on downstream task performance.