IC-791BLIP-2, BLIP, and COCA produce captions of comparable quality for zero-shot compositional image retrieval

Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep Akata

SourceVision-by-Language for Training-Free Compositional Image Retrieval

In an ablation on the CIRCO validation set, swapping the captioner in the CIR-EVL pipeline between BLIP-2 (map@5 13.33), BLIP (13.1), and COCA (13.64) yields only minor differences. The paper concludes that most state-of-the-art public captioning models perform similarly well, and retains BLIP-2 with a FLAN-T5 language model for generality.

Evidence
correlational
Key metric
CIRCO validation map@5: COCA 13.64, BLIP-2 13.33, BLIP 13.1
Caveat
Differences are small (0.54 points range); the comparison is on a single validation set and a single task.
Model
BLIP-2, BLIP, COCA
Datasets
CIRCO [eval]
Related findings
IC-788, IC-789, IC-790
Extraction
automatic-extraction