ViT

image · discriminative · anchor

Vision transformer applying a plain transformer to sequences of image patches. The backbone most of the vision-language models in the corpus build on.

Note
anchor found by search rather than in a citing paper, and checked against this entry's own description before it was recorded: "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale"
Variants
ViT-B/16, IN1K ViT, ViT-B, ViT-B/32, ViT-H, ViT-L/32, ViT/L-16, RViT-B/16, ViT-L, ViT-S, ViT-S/16, ViT-t/16

Findings

Shared mechanisms