IC-157GPT-4o's MediConfusion performance is robust to prompt format while InstructBLIP is highly sensitive and LLaVA-Med fails completely on multiple-choice evaluation

Mohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi Soltanolkotabi

SourceMediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models

An ablation study tests 11 manually crafted prompt variations on GPT-4o, InstructBLIP, and LLaVA-Med using multiple-choice evaluation. GPT-4o's set accuracy remains stable at 19.12% ± 0.04 across all prompts. InstructBLIP's set accuracy ranges from 1.25% to 12.05% depending on the prompt, with a mean of 4.79% ± 0.12, showing high sensitivity to phrasing. LLaVA-Med scores 0.00% on every prompt variation, indicating it cannot follow the multiple-choice format regardless of how the instruction is phrased. The authors note that LLaVA-Med was not specifically trained for multiple-choice QA.

Evidence
correlational
Key metric
GPT-4o: 19.12% ± 0.04 (all prompts). InstructBLIP: 4.79% ± 0.12 (range 1.25% to 12.05%). LLaVA-Med: 0.00% ± 0.00 (all prompts).
Caveat
Only 10 samples per prompt per model; the prompt variations were manually created by the authors, not generated by an LLM, to avoid LLM-generated text artifacts.
Model
GPT-4o, InstructBLIP, LLaVA-Med
Concepts
Failure mode
Related findings
IC-155, IC-156, IC-158
Extraction
automatic-extraction