SigLIP

Google DeepMind · 2023-03 · image, text · discriminative · anchor

Vision-language dual encoder trained with a pairwise sigmoid loss instead of the softmax contrastive loss used by CLIP, which removes the need for a global normalisation over the batch.

Variants
SigLIP ViT-B/16, SigLIP ViT-L/16

Findings

Shared mechanisms