IC-790GPT-4 outperforms GPT-3.5-turbo, Vicuna-13B, and Llama2-70B for generating target captions in zero-shot compositional image retrieval

Shyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep Akata

SourceVision-by-Language for Training-Free Compositional Image Retrieval

On the CIRCO validation set, replacing the LLM in the CIR-EVL pipeline with different models yields a clear performance hierarchy: GPT-4 achieves map@5 of 15.63, GPT-3.5-turbo 13.33, Vicuna-13B 12.65, Llama2-70B 10.22, and a template-based baseline (no LLM reasoning) 9.22. The paper concludes that textual reasoning by the LLM is critical and that stronger reasoning capabilities translate directly into better retrieval.

Evidence
correlational
Key metric
CIRCO validation map@5: GPT-4 15.63, GPT-3.5-turbo 13.33, Vicuna-13B 12.65, Llama2-70B 10.22, template (no LLM) 9.22
Caveat
Results are on the CIRCO validation set only; the task is specific to compositional image retrieval caption generation, not a general LLM benchmark.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, Vicuna Vicuna-13B, Llama 2 / Llama 2 base Llama 2 70B
Datasets
CIRCO [eval]
Related findings
IC-788, IC-789, IC-791
Extraction
automatic-extraction