The paper constructs a hallucination benchmark with six scenarios (natural selection, distraction, counterfactual reasoning, co-occurrence, misleading, OCR) each covering five tasks. Across all 18 evaluated models (8 text-to-image, 10 image-to-text), the average non-hallucination accuracy remains below 50%. The best text-to-image model (FLUX) averages 39.7% and the best image-to-text model (Nova Pro) averages 46.9%. Spatial reasoning is the most challenging task, with all text-to-image models achieving below 3% accuracy on natural prompts.
Evidence
correlational
Key metric
average performance for all mmfms in terms of non-hallucination accuracy is below 50%; best t2i model flux averages 39.7%, best i2t model nova pro averages 46.9%; spatial reasoning accuracy below 3% for all t2i models on natural prompts
Caveat
The benchmark data was selected to be challenging based on surrogate model performance, so the absolute numbers may not reflect performance on easier distributions.