IC-564The implicit attention matrices of Mamba, RWKV, and Griffin exhibit depth-dependent structure, with dependencies between distant tokens becoming more apparent in deeper layers

Itamar Zimerman, Ameen Ali Ali, Lior Wolf

SourceExplaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation

Using the paper's unified implicit attention formulation, the authors visualize the attention matrices of Mamba 2.8B, RWKV-430M, and Griffin at 25%, 50%, and 75% of layer depth. In all three architectures, the matrices show that long-range token dependencies are weak or absent in early layers and become progressively sharper in deeper layers. The authors note this echoes the earlier finding of Ali et al. (2024) for Mamba's S6 layer alone, but here it is confirmed for the full architecture including conv1d and gating. Additionally, RWKV matrices are characterized by distinct horizontal tiles while Mamba displays a more continuous structure.

Evidence
observational
Caveat
The observation is qualitative, based on visual inspection of attention matrix heatmaps at three depth fractions (25%, 50%, 75%) with a uniform prompt of size 32; no quantitative metric is reported for the depth-dependent trend.
Model
Mamba, RWKV, Griffin
Concepts
Depth-dependent structure
Related work
Ali et al. 2024 (The Hidden Attention of Mamba Models) [builds-on]
Extraction
automatic-extraction