IC-847CLIP ViT-B/32 CLIPScore achieves only ρ=0.276 / τ=0.191 correlation with human 1-5 likert T2I alignment ratings on TIFA160

Jaemin Cho, Yushi Hu, Jason Michael Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, Su Wang

SourceDavidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

The paper reports the Spearman and Kendall correlation between CLIPScore (computed with CLIP ViT-B/32) and human 1-5 likert text-image consistency ratings on TIFA160 prompts across five T2I models. The correlation is substantially lower than the best QG/A combination (DSG+Pali: ρ=0.571, τ=0.458), supporting the paper's argument that single-summary embedding similarity is a weaker proxy for fine-grained T2I alignment than question-answering-based evaluation.

Evidence
correlational
Key metric
CLIPScore: ρ=0.276, τ=0.191; compared to DSG+Pali: ρ=0.571, τ=0.458
Caveat
The human ratings were collected by different raters than those in the original TIFA paper, so direct comparison with TIFA's reported CLIPScore τ=23.1 is not valid.
Model
CLIP / CLIP-ViT (LC)
Methods
CLIPScore [primary], Spearman's ρ [eval], Kendall's τ [eval]
Related findings
IC-848
Extraction
automatic-extraction