IC-364Llama-3-8B hidden states are zero-mean unimodal, with gaussian-like distributions before attention and MLP blocks and laplacian-like distributions in intermediate states

James Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo, Yoon Kim, Ben Athiwaratkun

SourceTraining-Free Activation Sparsity in Large Language Models

The paper collects activations from Llama-3-8B sampled from C4 at four hidden states within transformer blocks 8, 16, and 24. The hidden states preceding the attention and MLP layers exhibit gaussian-like shapes, while the intermediate states within those blocks exhibit laplacian-like shapes. All distributions are zero-mean and unimodal, with the concentration around zero motivating magnitude-based pruning. The authors note that some activations are heavy-tailed and contain outliers, consistent with prior work.

Evidence
observational
Caveat
The authors explicitly state they do not attempt to explain why these distributions are shaped this way, nor give theoretical underpinnings for why activation sparsity works.
Model
Llama 3 8B
Datasets
C4 [source]
Related work
DejaVu [context]
Related findings
IC-365, IC-366, IC-367
Extraction
automatic-extraction