Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
MSR-VTT
anchor
Findings
IC-390
The parallelotope volume of modality embeddings from LanguageBind, VAST, and Valor on MSR-VTT is strongly correlated with their downstream R@1 retrieval performance
[eval]
IC-981
ImageBind's indirect alignment through images degrades zero-shot performance on non-visual modalities and prevents emergent cross-modal retrieval
[eval]