In Table 2, FIXLIP explanations of SigLIP-2 obtain higher pointing-game recognition than CLIP at both reported patch configurations, ViT-B/32 and ViT-B/16. The authors interpret this as SigLIP-2 being more faithfully explainable through cross-modal interactions and suggest that smaller models may be more faithfully explainable than larger ones.
Evidence
correlational
Key metric
pointing game recognition for 1 to 4 objects, ViT-B/32: SigLIP-2 0.90 / 0.90 / 0.89 / 0.89 against CLIP 0.83 / 0.81 / 0.82 / 0.85; ViT-B/16: SigLIP-2 0.86 / 0.88 / 0.87 / 0.87 against CLIP 0.81 / 0.81 / 0.81 / 0.82
Caveat
The comparison does not use exactly the same pointing-game cases: subgames containing goldfish or ipod are omitted for SigLIP models because their tokenizer splits those labels into multiple tokens. The authors offer no mechanism for the SigLIP-2 result. Their scale interpretation is not uniform across families: SigLIP scores no lower at ViT-L/16 than at ViT-B/16.