Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Locality Alignment Improves Vision-Language Models
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-499
ViT patch embeddings contain local semantic information beyond the [cls] token, as shown by performance degradation when restricting the output head to [cls] only or removing positional embeddings