IC-127ImageBind's direct evaluation closely matches logsumexp for both vision-language and audio-language alignment, validating the law for ImageBind

Yongwei Che, Benjamin Eysenbach

SourceThe "Law'' of the Unconscious Contrastive Learner: Probabilistic Alignment of Unpaired Modalities

The paper tests ImageBind (which aligns modalities through a shared image space) in two configurations: vision-language alignment via audio, and audio-language alignment via images. In both cases, the logsumexp method's recall@1 closely matches ImageBind's own direct evaluation. For vision-language, logsumexp achieves 31.4% ± 1.7% versus ImageBind's direct 32.4% ± 1.5%. For audio-language, logsumexp achieves 29.4% ± 3.3% versus ImageBind's direct 29.1% ± 1.7%. The near-identical performance confirms that the inner product between unpaired modality representations recovers the correct similarity for ImageBind.

Evidence
correlational
Key metric
Vision-language: 31.4% ± 1.7% (logsumexp) vs 32.4% ± 1.5% (direct ImageBind); Audio-language: 29.4% ± 3.3% (logsumexp) vs 29.1% ± 1.7% (direct ImageBind)
Caveat
Results are over 100 trials with 95% confidence intervals; recall@1 is evaluated from a set of 25 samples.
Model
ImageBind
Concepts
Linear representation, Distance preservation
Datasets
AudioSet [eval]
Related findings
IC-124, IC-125, IC-126
Extraction
automatic-extraction