IC-495All evaluated multimodal foundation models achieve average non-hallucination accuracy below 50% across six hallucination scenarios

Chejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie, Mintong Kang, Yujin Potter, Zhun Wang, Zhuowen Yuan, Alexander Xiong, Zidi Xiong, Chenhui Zhang, Lingzhi Yuan, Yi Zeng, Peiyang Xu, Chengquan Guo, Andy Zhou, Jeffrey Ziwei Tan, Xuandong Zhao, Francesco Pinto, Zhen Xiang, Yu Gai, Zinan Lin, Dan Hendrycks, Bo Li, Dawn Song

SourceMMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

The paper constructs a hallucination benchmark with six scenarios (natural selection, distraction, counterfactual reasoning, co-occurrence, misleading, OCR) each covering five tasks. Across all 18 evaluated models (8 text-to-image, 10 image-to-text), the average non-hallucination accuracy remains below 50%. The best text-to-image model (FLUX) averages 39.7% and the best image-to-text model (Nova Pro) averages 46.9%. Spatial reasoning is the most challenging task, with all text-to-image models achieving below 3% accuracy on natural prompts.

Evidence
correlational
Key metric
average performance for all mmfms in terms of non-hallucination accuracy is below 50%; best t2i model flux averages 39.7%, best i2t model nova pro averages 46.9%; spatial reasoning accuracy below 3% for all t2i models on natural prompts
Caveat
The benchmark data was selected to be challenging based on surrogate model performance, so the absolute numbers may not reflect performance on easier distributions.
Model
FLUX / FLUX1, DALL·E 3, DALL·E 2, Stable Diffusion SDXL, Nova Pro, GPT-4o, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4V / GPT-4 vision, Llama-3-2-Vision Llama-3.2-90B-Vision-Instruct, LLaVA-NeXT / LLaVA 1.6, Gemini 1.5 / Gemini Pro 1.5
Concepts
Failure mode
Datasets
MS COCO / COCO / COCO 2014 / COCO 2017 / COCO 20k / COCO-it / COCO-wl [source], HRS-Bench [source]
Methods
GroundingDINO [eval], LLaVA-Next [eval]
Related findings
IC-496, IC-497, IC-498
Extraction
automatic-extraction