IC-491Presenting three images as a single collage rather than sequentially reduces MLLM accuracy by approximately 40% on the relationship application step
Nilay Yilmaz, Maitreya Patel, Yiran Lawrence Luo, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang
GPT-4o and Qwen2-VL-7B-Instruct, which accept both input formats, were evaluated on VOILA-WD and VOILA-ND using L2M prompting. When the three input images are combined into a single collage, accuracy on the applying-relationship step drops substantially: GPT-4o goes from 6.44% to 3.94% on VOILA-WD and from 29.03% to 19.43% on VOILA-ND. The authors attribute the degradation to image resolution constraints in the collage format. A follow-up test with LLaVA-OneVision's anyres strategy partially closes the gap (53% vs. 57% on image description).
Evidence
correlational
Key metric
GPT-4o applying relationship: 3.94% (collage) vs. 6.44% (sequential) on VOILA-WD; 19.43% vs. 29.03% on VOILA-ND; Qwen2-VL: 0.52% vs. 0.85% (VOILA-WD); 3.77% vs. 6.8% (VOILA-ND)
Caveat
Only two models (GPT-4o, Qwen2-VL) support both input formats; the effect may be specific to models with resolution constraints on multi-image input.