IC-1429OPT 6.7B exhibits aggregated sparsity where approximately 50% of neurons remain unused across the first 150 tokens, with a non-random reuse pattern enabling 1.27x speculative decoding speedup at gamma=16

Seyed Iman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, Mehrdad Farajtabar

SourceReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

The paper measures how many distinct neurons are activated across a sequence of generated tokens in OPT 6.7B. On average, about 50% of all FFN neurons are never used across the first 150 tokens of WikiText prompts. This aggregated sparsity is significantly higher than what would be expected from random per-token activation (the observed curve exceeds the random sparsity curve in Fig 7b), indicating a structured reuse of neurons. Exploiting this temporal pattern, the paper shows that reusing previously loaded weights for subsequent tokens causes only a slight perplexity increase, and when integrated into speculative decoding, yields a 1.27x speedup at gamma=16 versus 1.20x for random sparsity.

Evidence
observational
Key metric
~50% neurons unused across first 150 tokens (WikiText); speculative decoding speedup 1.27x (aggregated) vs 1.20x (random) at gamma=16; at gamma=64, aggregated sparsity speedup ~1.14x vs negligible for random
Caveat
Measured on WikiText prompts; the paper notes results hold for other ReLU models and datasets but does not quantify the variation
Model
OPT 6.7B
Datasets
WikiText [eval]
Methods
Speculative Decoding [primary]
Related work
Speculative Decoding [builds-on]
Related findings
IC-1428
Extraction
automatic-extraction