Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
2025-01-22
· ICLR 2025 Poster ·
anchor
Findings
IC-564
The implicit attention matrices of Mamba, RWKV, and Griffin exhibit depth-dependent structure, with dependencies between distant tokens becoming more apparent in deeper layers