IC-600Chain-of-thought prompting improves most LMMs on synthetic detection but degrades LLaVA-ov-7b from 56.6% to 18.8%, while GPT-4o performs well without it (64.1% baseline)
Junyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang, Tianyi Bai, Hengrui Kang, Jun He, Honglin Lin, Zihao Wang, Tong Wu, Zhizheng Wu, Yiping Chen, Dahua Lin, Conghui He, Weijia Li
On image and 3D judgment tasks, CoT prompting raises accuracy for InternVL2-8B (49.6→50.4), Qwen2-VL-7B (56.8→59.5), Gemini-1.5-Pro (47.9→51.0), and Claude-3.5-Sonnet (55.2→56.4). GPT-4o already scores 64.1% without CoT and reaches 75.1% with it. Few-shot prompting fails to help most models. LLaVA-ov-7b suffers a dramatic drop from 56.6% to 18.8% under CoT, which the paper attributes to degraded long-context understanding after fine-tuning.
The paper notes that the CoT degradation in LLaVA-ov-7b may result from a decline in long-context ability after fine-tuning, and that more CoT results are in the appendix.