Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
JoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and Attention
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-913
In OPT-2.7B, Pythia-70M/1.4B/6.9B, and BERT-base, the stable rank of MLP lower layers shows a drop-and-bounce pattern during training that is more salient in top layers while bottom layers show suppressed dropping curves
IC-914
In Pythia models (70M through 2.8B), BERT-base, OPT-6.7B, LLaMA-2-7B, and ViT-Huge, the MLP out-projection vectors are almost orthogonal throughout training
IC-915
In Pythia-70M and Pythia-160M, individual MLP hidden neurons are activated by multiple irrelevant token combinations (pattern superposition)