IC-157GPT-4o's MediConfusion performance is robust to prompt format while InstructBLIP is highly sensitive and LLaVA-Med fails completely on multiple-choice evaluation
Mohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi Soltanolkotabi
An ablation study tests 11 manually crafted prompt variations on GPT-4o, InstructBLIP, and LLaVA-Med using multiple-choice evaluation. GPT-4o's set accuracy remains stable at 19.12% ± 0.04 across all prompts. InstructBLIP's set accuracy ranges from 1.25% to 12.05% depending on the prompt, with a mean of 4.79% ± 0.12, showing high sensitivity to phrasing. LLaVA-Med scores 0.00% on every prompt variation, indicating it cannot follow the multiple-choice format regardless of how the instruction is phrased. The authors note that LLaVA-Med was not specifically trained for multiple-choice QA.
Only 10 samples per prompt per model; the prompt variations were manually created by the authors, not generated by an LLM, to avoid LLM-generated text artifacts.