Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Skip-Attention: Improving Vision Transformers by Paying Less Attention
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1540
CLS-token attention maps in pretrained ViT-t/16 exhibit high inter-layer correlation (cosine similarity up to 0.97) concentrated in layers 3–10, and MSA block outputs show high CKA in layers 2–8