Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-250
Existing multimodal embedding models show highly uneven performance across MMEB's four meta-task categories, with VQA scores as low as 4.2 and overall scores ranging from 13.3 to 44.7
IC-251
CLIP's overall MMEB performance drops by 29.4% when task-specific instructions are prepended to queries, with classification degrading by 59.3%