IC-480Pre-trained M-LLMs (GPT-4o, LLaVA-v1.6-34B, InternVL2-26B, Qwen2-VL-7B) produce imprecise tampering explanations when artifacts require fine-grained pixel-level analysis such as lighting or perspective inconsistencies

Zhipei Xu, Xuanyu Zhang, Runyi Li, Zecheng Tang, Qing Huang, Jian Zhang

SourceFakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models

The paper evaluates four released M-LLMs on their ability to generate textual explanations of image tampering, measured by cosine semantic similarity (CSS) against GPT-4o-generated ground-truth descriptions across nine datasets spanning photoshop, deepfake, and AIGC-editing. All four models achieve CSS scores between 0.4887 and 0.6760, substantially below the fine-tuned FakeShield (0.7537–0.8873). The authors note that these M-LLMs can leverage pre-training knowledge to make reasonable judgments when tampering causes clear physical-law violations, but they struggle with more precise analyses like detecting lighting or perspective inconsistencies, which reduces overall explanation accuracy.

Evidence
correlational
Key metric
CSS scores: GPT-4o 0.5183–0.6289, LLaVA-v1.6-34B 0.5034–0.6457, InternVL2-26B 0.5750–0.6760, Qwen2-VL-7B 0.4887–0.6209 across CASIA1+, IMD2020, Columbia, Coverage, NIST, DSO, Korus, deepfake, AIGC-editing; FakeShield 0.7537–0.8873
Caveat
Ground-truth descriptions were generated by GPT-4o itself, so GPT-4o's CSS is a self-consistency measure rather than an independent reference; the paper does not report per-artifact-type breakdowns of where each model fails.
Model
GPT-4o, LLaVA-NeXT / LLaVA 1.6 LLaVA-v1.6-34B, InternVL2 InternVL2-26B, Qwen2-VL Qwen2-VL-7B
Concepts
Failure mode
Datasets
CASIA1+ [eval], IMD2020 [eval], DFFD [eval]
Methods
Cosine similarity / Cosine similarity analysis / Cosine semantic similarity / cosine similarity of hidden states / Sample-wise cosine similarity / Cosine similarity of attention maps / Cosine similarity perturbation analysis / Cosine similarity template matching / Cosine similarity to neighbours / Semantic consistency (cosine similarity) [eval]
Extraction
automatic-extraction