IC-126CLIP and CLAP language representations are statistically indistinguishable from a uniform distribution on the hypersphere

Yongwei Che, Benjamin Eysenbach

SourceThe "Law'' of the Unconscious Contrastive Learner: Probabilistic Alignment of Unpaired Modalities

The paper tests assumption 3 (that contrastive representations are uniform on the unit hypersphere) on real-world data. It computes language representations for each model over the AudioSet ontology and performs a two-sample Kolmogorov-Smirnov test against a uniform hypersphere distribution. CLIP yields a p-value of 0.0877 and CLAP a p-value of 0.1788, both above the conventional 0.05 significance threshold, so neither model's representations show significant deviation from uniformity.

Evidence
correlational
Key metric
KS test p-value 0.0877 (CLIP), p-value 0.1788 (CLAP) against uniform hypersphere
Caveat
The test is performed on language representations over the AudioSet ontology only, not on image or audio representations.
Model
CLIP / CLIP-ViT (LC), CLAP
Datasets
AudioSet [eval]
Methods
Kolmogorov-Smirnov two-sample test [primary]
Related findings
IC-124, IC-125, IC-127
Extraction
automatic-extraction