IC-156Gemini models show substantially lower confusion scores than other MLLMs yet still perform at or near random guessing, suggesting their bottleneck is medical knowledge or reasoning rather than visual encoding

Mohammad Shahab Sepehri, Zalan Fabian, Maryam Soltanolkotabi, Mahdi Soltanolkotabi

SourceMediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models

Gemini 1.5 Pro and Gemini 2.0 Flash have confusion scores of 58.52% and 67.05% respectively, far below the 80-100% seen in most other models. This means they do differentiate between the two images in a pair more often than other models. However, their set accuracies (19.89% and 29.55%) are still at or only slightly above the 25% random baseline. The authors hypothesize that Gemini's visual representations are rich enough to distinguish the images, but the models lack the medical knowledge or reasoning skills to map that distinction to the correct answer. Gemini 2.0 Flash achieves the best individual accuracy (61.93%) and the best category-specific result (42.86% set accuracy on nuclear medicine).

Evidence
correlational
Key metric
Gemini 1.5 Pro: confusion 58.52%, set acc 19.89%, indiv acc 51.14%. Gemini 2.0 Flash: confusion 67.05%, set acc 29.55%, indiv acc 61.93%. Compare LLaVA-Med confusion 97.16%, RadFM confusion 85.80%.
Caveat
The authors note this is a hypothesis; they do not directly measure medical knowledge or reasoning ability separately from visual encoding.
Model
Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Gemini 2.0 Flash
Concepts
Failure mode
Related findings
IC-155, IC-157, IC-158
Extraction
automatic-extraction