IC-1428OPT 6.7B exhibits over 90% activation sparsity in FFN layers, reducing inference from 6.6G to 4.5G flops per token, while Llama 7B (SiLU) and Falcon 7B (GELU) show near-zero sparsity

Seyed Iman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, Mehrdad Farajtabar

SourceReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

The paper measures the fraction of zero activations in the feed-forward network between up and down projections across all layers of three released LLMs. OPT 6.7B, which natively uses ReLU, shows sparsity exceeding 90% in every layer, zeroing out 95% of rows in the down-projection weight matrix. This structural sparsity reduces the FLOPs per token from 6.6G to 4.5G, a 32% computation saving. In contrast, Llama 7B (SiLU) and Falcon 7B (GELU) show near-zero activation sparsity in the same measurement, confirming that the sparsity is a direct consequence of the ReLU activation function.

Evidence
observational
Key metric
OPT 6.7B: all layers >90% sparsity, 95% of down-projection rows zeroed, 6.6G to 4.5G flops per token (32% saving); Llama 7B and Falcon 7B: near-zero sparsity (Fig 1a)
Model
OPT 6.7B, LLaMA Llama 7B, Falcon Falcon-7B
Datasets
WikiText [eval]
Related work
Sparse is Enough in Scaling Transformers [context]
Related findings
IC-1429
Extraction
automatic-extraction