SourceADIFF: Explaining audio difference using natural language
The paper introduces the audio difference explanation task and evaluates Qwen-Audio 7B in zero-shot mode (no task-specific training) on human-rated correctness, granularity, and readability across three scenarios: studio recordings, FSD50K samples, and GTZAN music. Qwen-Audio zero-shot achieves average scores of 2.81, 2.56, and 3.23 out of 5, respectively, which are lower than the authors' 128M parameter naive baseline (3.26, 3.35, 3.39) and the full ADIFF model (3.47, 3.53, 3.57). Despite being a 7B parameter SOTA audio-language model and the only ALM in the literature supporting two audio inputs, Qwen-Audio zero-shot does not leverage its scale advantage on this comparative reasoning task without task-specific training.