Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Memorization Capacity of Multi-Head Attention in Transformers
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-935
In pre-trained ViT, query vector Kruskal rank reaches the context size only after one self-attention layer, while general position fails at all depths
IC-936
GPT-2's learned positional encodings cause context vectors to lose linear independence after one layer, whereas BERT's sinusoidal encodings preserve it