Light Dark Google DeepMind · 2023-03 · image, text · discriminative · anchor
Vision-language dual encoder trained with a pairwise sigmoid loss instead of the softmax contrastive loss used by CLIP, which removes the need for a global normalisation over the batch.
Variants SigLIP ViT-B/16 , SigLIP ViT-L/16 Findings Shared mechanisms Distance preservation also in CLIP / CLIP-ViT (LC) , CoPlace , DINO , DINOv2 , Gemma , Gemma 2 , ImageBind , LanguageBind , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , Llama-3.2-3B , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-Med , MAE , MAE-B/16 , OpenCLIP , Pythia , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SLIP , TerraMind , ViT Explanation faithfulness also in BakLLaVA , CF2 , Claude 3 , Claude 3.5 , CLIP / CLIP-ViT (LC) , DRUM , Fuyu , GEM , Gemini 1.5 / Gemini Pro 1.5 , Gemma 2 , GPT-3.5 / ChatGPT-3.5 , GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report , GPT-4o , GPT-J , Idefics , Llama 3.1 , LLaVA-NeXT / LLaVA 1.6 , MAE-B/16 , MobileNetV2 , mPLUG-Owl3 , OpenFlamingo , PGExplainer , Pythia , RCExplainer , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SigLIP-2 , Stable Diffusion , TAGExplainer , ViT Linear representation also in BLOOM , Cambrian-1 , Chameleon , CLIP / CLIP-ViT (LC) , DINOv2 , EVA-CLIP , Falcon , Gemma , Gemma 2 , GPT-2 , GPT-J , HPSv2 , ImageBind , InstructBLIP , LanguageBind , LLaMA , Llama 2 / Llama 2 base , Llama 3 , Llama 3.1 , Llama-3.2-3B , LLaVA-1.5 / LLaVA-v1.5 , LLaVA-NeXT / LLaVA 1.6 , MAE , Mamba , Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1 , OLMo / OLMo base , OpenCLIP , Phi-3 , PickScore , Pythia , Qwen2-VL , Qwen2.5 , ResNet / ResNet-152 / ResNet-101 / ResNet-50-BN , SALMONN , SAM , TerraMind , Tulu 2 , Vicuna , ViT