In an ablation on the CIRCO validation set, swapping the captioner in the CIR-EVL pipeline between BLIP-2 (map@5 13.33), BLIP (13.1), and COCA (13.64) yields only minor differences. The paper concludes that most state-of-the-art public captioning models perform similarly well, and retains BLIP-2 with a FLAN-T5 language model for generality.