Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
AudioSet
anchor
Findings
IC-124
Direct comparison of CLIP image embeddings with CLAP audio embeddings achieves near-chance retrieval, while logsumexp bridging through the shared language modality recovers 62% recall@10 on AudioSet
[eval]
IC-125
LanguageBind's direct evaluation achieves 70% recall@10 on AudioSet, empirically validating that the inner product between unpaired modality representations recovers the correct probability ratio
[eval]
IC-126
CLIP and CLAP language representations are statistically indistinguishable from a uniform distribution on the hypersphere
[eval]
IC-127
ImageBind's direct evaluation closely matches logsumexp for both vision-language and audio-language alignment, validating the law for ImageBind
[eval]
IC-981
ImageBind's indirect alignment through images degrades zero-shot performance on non-visual modalities and prevents emergent cross-modal retrieval
[eval]