Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Vision-by-Language for Training-Free Compositional Image Retrieval
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-788
CLIP ViT-B/32 fails to retrieve the correct image even when the generated target caption is well-aligned with the ground-truth image
IC-789
CLIP retrieval quality in zero-shot compositional image retrieval scales log-linearly with model size from approximately 150M to 2.5B parameters
IC-790
GPT-4 outperforms GPT-3.5-turbo, Vicuna-13B, and Llama2-70B for generating target captions in zero-shot compositional image retrieval
IC-791
BLIP-2, BLIP, and COCA produce captions of comparable quality for zero-shot compositional image retrieval