Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Training-Free Activation Sparsity in Large Language Models
2025-01-22
· ICLR 2025 Spotlight ·
anchor
Findings
IC-364
Llama-3-8B hidden states are zero-mean unimodal, with gaussian-like distributions before attention and MLP blocks and laplacian-like distributions in intermediate states
IC-365
70B LLM variants tolerate substantially higher activation sparsity than smaller counterparts, and Llama-3 shows more degradation than Llama-2 and Mistral at 50% sparsity
IC-366
In Llama-3-70B, activation sparsifiability varies systematically across depth: Wq/Wk peak in block 0 then decline sharply, Wo peaks at 80-90% mid-model, and Wdown is consistently more sparsifiable than Wgate and Wup
IC-367
Sparsifying initial tokens of the prefill phase causes disproportionate degradation in Llama-3-8B due to attention sink behavior