IC-600Chain-of-thought prompting improves most LMMs on synthetic detection but degrades LLaVA-ov-7b from 56.6% to 18.8%, while GPT-4o performs well without it (64.1% baseline)

Junyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang, Tianyi Bai, Hengrui Kang, Jun He, Honglin Lin, Zihao Wang, Tong Wu, Zhizheng Wu, Yiping Chen, Dahua Lin, Conghui He, Weijia Li

SourceLOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

On image and 3D judgment tasks, CoT prompting raises accuracy for InternVL2-8B (49.6→50.4), Qwen2-VL-7B (56.8→59.5), Gemini-1.5-Pro (47.9→51.0), and Claude-3.5-Sonnet (55.2→56.4). GPT-4o already scores 64.1% without CoT and reaches 75.1% with it. Few-shot prompting fails to help most models. LLaVA-ov-7b suffers a dramatic drop from 56.6% to 18.8% under CoT, which the paper attributes to degraded long-context understanding after fine-tuning.

Evidence
correlational
Key metric
Image & 3D judgment: GPT-4o baseline 64.1, CoT 75.1, few-shot 74.2; LLaVA-ov-7b baseline 56.6, CoT 18.8, few-shot 46.4; InternVL2-8b baseline 49.6, CoT 50.4; Qwen2-VL-7b baseline 56.8, CoT 59.5 (Table 5)
Caveat
The paper notes that the CoT degradation in LLaVA-ov-7b may result from a decline in long-context ability after fine-tuning, and that more CoT results are in the appendix.
Model
GPT-4o, LLaVA-OneVision LLaVA-ov-7B, InternVL2 InternVL2-8B, Qwen2-VL Qwen2-VL-7B, Gemini 1.5 / Gemini Pro 1.5 Gemini 1.5 Pro, Claude 3.5 Sonnet
Methods
Chain-of-Thought prompting / Chain-of-Thought (CoT) prompting / CoT prompting / Few-shot Chain of Thought / Few-shot CoT / Chain-of-Thought (CoT-S) / CoT-bag prompting / Few-shot CoT prompting / Rationale prompting / Zero-shot CoT prompting / Wei et al. (2022) chain-of-thought prompting [primary], Few-shot prompting / 2-shot prompting / Few-shot ICL / Few-shot prompting for base models [compared-to]
Related findings
IC-597, IC-598, IC-599
Extraction
automatic-extraction