IC-788CLIP ViT-B/32 fails to retrieve the correct image even when the generated target caption is well-aligned with the ground-truth image

Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep Akata

SourceVision-by-Language for Training-Free Compositional Image Retrieval

Using TIFA to measure text-image alignment, the paper shows that the modified caption produced by the LLM aligns well with the ground-truth target image, yet CLIP ViT-B/32 retrieves a different image whose actual alignment with the caption is notably lower. The paper concludes that CLIP retrieval is a severe bottleneck: the captioning and reasoning stages produce valid target descriptions that the retrieval model cannot match. This was observed across the CIRCO validation set.

Evidence
correlational
Caveat
The TIFA scores are shown in a figure (Fig. 5) without explicit numeric values printed in the text; the claim is qualitative ('notably lower').
Model
CLIP / CLIP-ViT (LC)
Concepts
Failure mode
Datasets
CIRCO [eval]
Methods
TIFA [eval]
Related work
SugarCrepe [context]
Related findings
IC-789, IC-790, IC-791
Extraction
automatic-extraction