Six audio-capable LMMs (Qwen-Audio, Salmonn-7B, AnyGPT, OneLLM, LTU, Gemini-1.5-Flash) were tested on 2280 audio judgment questions spanning speech, singing, environmental sound, and music. All perform at or below 51.2% accuracy, indistinguishable from the 50% random baseline. The paper attributes this to audio LMMs being trained for content comprehension rather than acoustic feature sensitivity, and notes that even the best-performing expert model AASIST (69.4%) only matches human level.
The paper notes that few open-source and proprietary LMMs support the audio modality, limiting the generalizability of this finding to the broader LMM landscape.