Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
2024-01-16
· ICLR 2024 oral ·
anchor
Findings
IC-1428
OPT 6.7B exhibits over 90% activation sparsity in FFN layers, reducing inference from 6.6G to 4.5G flops per token, while Llama 7B (SiLU) and Falcon 7B (GELU) show near-zero sparsity
IC-1429
OPT 6.7B exhibits aggregated sparsity where approximately 50% of neurons remain unused across the first 150 tokens, with a non-random reuse pattern enabling 1.27x speculative decoding speedup at gamma=16