CLIP / CLIP-ViT (LC)

OpenAI · 2021-02-26 · image, text · discriminative · anchor · artifact

Dual-encoder model that maps images and text into a shared space and scores their similarity.

Variants
CLIP RN50, CLIP RN101, CLIP RN50x4, CLIP ViT-B/32, CLIP ViT-B/16, CLIP ViT-L/14, CLIP ViT-L/14@336px, CLIP-B/32

Findings

Shared mechanisms