IC-599All evaluated audio LMMs perform at or near random chance (44.4%–51.2%) on synthetic audio detection, while humans achieve 69.2%

Junyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang, Tianyi Bai, Hengrui Kang, Jun He, Honglin Lin, Zihao Wang, Tong Wu, Zhizheng Wu, Yiping Chen, Dahua Lin, Conghui He, Weijia Li

SourceLOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

Six audio-capable LMMs (Qwen-Audio, Salmonn-7B, AnyGPT, OneLLM, LTU, Gemini-1.5-Flash) were tested on 2280 audio judgment questions spanning speech, singing, environmental sound, and music. All perform at or below 51.2% accuracy, indistinguishable from the 50% random baseline. The paper attributes this to audio LMMs being trained for content comprehension rather than acoustic feature sensitivity, and notes that even the best-performing expert model AASIST (69.4%) only matches human level.

Evidence
correlational
Key metric
Audio judgment accuracy: Qwen-Audio 49.8%, Salmonn-7B 51.2%, AnyGPT 49.8%, OneLLM 49.9%, LTU 44.4%, Gemini-1.5-Flash 49.4%; human 69.2%; AASIST 69.4% (Table 16)
Caveat
The paper notes that few open-source and proprietary LMMs support the audio modality, limiting the generalizability of this finding to the broader LMM landscape.
Model
Qwen-Audio, SALMONN SALMONN-7B, OneLLM, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Flash, AASIST
Concepts
Failure mode
Related findings
IC-597, IC-598, IC-600
Extraction
automatic-extraction