SourceExplaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
Using the paper's unified implicit attention formulation, the authors visualize the attention matrices of Mamba 2.8B, RWKV-430M, and Griffin at 25%, 50%, and 75% of layer depth. In all three architectures, the matrices show that long-range token dependencies are weak or absent in early layers and become progressively sharper in deeper layers. The authors note this echoes the earlier finding of Ali et al. (2024) for Mamba's S6 layer alone, but here it is confirmed for the full architecture including conv1d and gating. Additionally, RWKV matrices are characterized by distinct horizontal tiles while Mamba displays a more continuous structure.