IC-786CLIP-ViT (LC) achieves 0.87 accuracy and 0.91 average precision on fake image detection

Chunsan Hong, ByungHee Cha, Tae-Hyun Oh

SourceCAS: A Probability-Based Approach for Universal Condition Alignment Score

The authors trained CLIP-ViT (LC) on a small custom dataset of 100 real and 100 generated images per model (from DP2.0, RV1.4, Epic, SD2.1) and evaluated on 200 test images. CLIP-ViT (LC) achieved 0.87 accuracy and 0.91 average precision. A CAS ensemble MLP trained on the same data achieved 0.92 accuracy and 0.96 average precision, outperforming the state-of-the-art fake detector.

Evidence
correlational
Key metric
Table 7: CLIP-ViT (LC) accuracy 0.87, average precision 0.91; CAS ensemble MLP accuracy 0.92, average precision 0.96
Caveat
Small-scale experiment with only 500 training samples; the CAS ensemble uses four different diffusion models as features; the authors note the test and training generation models differ.
Model
CLIP / CLIP-ViT (LC)
Related findings
IC-783, IC-784, IC-785
Extraction
automatic-extraction