The paper introduces the Normalized Bias Index (NBI) to quantify whether a model preferentially classifies inputs as real or synthetic. A heatmap across 20+ models shows that most LMMs have a pronounced bias in at least one modality. GPT-4o specifically tends to classify textual data as real while being biased toward judging 3D data as AI-generated. Despite using two question phrasings to minimize cueing effects, the bias persists across the majority of evaluated models.
Evidence
correlational
Key metric
NBI values shown in Figure 5(a) heatmap; text states 'gpt-4o tends to classify textual data as real, whereas it is biased towards judging 3d data as ai-generated' and 'a pronounced bias is still evident across most models'
Caveat
The NBI values are presented only as a color-coded heatmap in Figure 5(a); no per-model numeric NBI values are printed in the text or tables.