SourceVision-by-Language for Training-Free Compositional Image Retrieval
Using TIFA to measure text-image alignment, the paper shows that the modified caption produced by the LLM aligns well with the ground-truth target image, yet CLIP ViT-B/32 retrieves a different image whose actual alignment with the caption is notably lower. The paper concludes that CLIP retrieval is a severe bottleneck: the captioning and reasoning stages produce valid target descriptions that the retrieval model cannot match. This was observed across the CIRCO validation set.