IC-598GPT-4o's synthetic image detection accuracy drops sharply on specialized domains (satellite 45.0%, medical 54.3%) compared to common image types (object 84.4%, person 84.4%)

Junyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang, Tianyi Bai, Hengrui Kang, Jun He, Honglin Lin, Zihao Wang, Tong Wu, Zhizheng Wu, Yiping Chen, Dahua Lin, Conghui He, Weijia Li

SourceLOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

Breaking down GPT-4o's image judgment accuracy by subcategory reveals a stark gap between common and specialized image types. On objects and persons, GPT-4o exceeds 84%, but on satellite imagery it falls to 45.0% and on medical images to 54.3%. The paper attributes this to a lack of expert domain knowledge in current LMMs, noting that performance on less-trained types like documents (60.1%) is also below the overall average of 63.4%.

Evidence
correlational
Key metric
GPT-4o judgment accuracy by image subcategory: overall 63.4%, scene 70.1%, animal 69.7%, person 84.4%, object 84.4%, medicine 54.3%, doc 60.1%, satellite 45.0% (Table 10)
Caveat
The paper notes that qwen2-vl-72b surpasses GPT-4o in some specialized image categories, suggesting the gap is model-specific rather than universal.
Model
GPT-4o, Qwen2-VL Qwen2-VL-72B
Concepts
Failure mode
Related findings
IC-597, IC-599, IC-600
Extraction
automatic-extraction