SourceJoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and Attention
The paper examines public intermediate checkpoints of OPT-2.7B and Pythia-70M/1.4B/6.9B to verify its theoretical prediction that attention first becomes sparse then dense. While the attention entropy patterns show only a less salient drop-and-bounce, the stable rank of the MLP lower-layer projection matrix shows a much more pronounced drop-and-bounce structure in top layers. In bottom layers, the stable rank shows only dropping curves, suppressed by top-level learning. The same pattern is confirmed in BERT-base (encoder-decoder), where attention entropy and stable rank behave similarly to the decoder-only case.