IC-981ImageBind's indirect alignment through images degrades zero-shot performance on non-visual modalities and prevents emergent cross-modal retrieval

Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, WANG HongFa, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Cai Wan Zhang, Zhifeng Li, Wei Liu, Li Yuan

SourceLanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

The paper measures ImageBind (ViT-Huge) on zero-shot classification across video, infrared, depth, and audio benchmarks, as well as on emergent cross-modal retrieval. ImageBind aligns non-visual modalities to images rather than directly to language, and this indirect alignment yields substantially lower accuracy on non-visual tasks: 63.4% top-1 on LLVIP, 54.0% on NYU-D, and 66.9% on ESC-50. Furthermore, ImageBind is unable to perform emergent zero-shot retrieval between non-visual modalities (video-to-audio, RGB-to-infrared, RGB-to-depth), marked as not achievable in the paper's Table 7, whereas direct language alignment enables such transfers.

Evidence
correlational
Key metric
top-1 accuracy (ImageBind huge): LLVIP 63.4%, NYU-D 54.0%, ESC-50 66.9%, AudioSet 17.6%, VGGSound 27.8%, K400 50.0%; emergent cross-modal retrieval (Table 7): not achievable (marked ✗) on AVE, VGGSound, LLVIP, NYU
Model
ImageBind
Concepts
Failure mode
Datasets
LLVIP [eval], NYU-v2 / NYU-D [eval], ESC-50 [eval], AudioSet [eval], VGGSound [eval], MSR-VTT [eval], AVE [eval]
Extraction
automatic-extraction