SourceTraining-Free Activation Sparsity in Large Language Models
The paper collects activations from Llama-3-8B sampled from C4 at four hidden states within transformer blocks 8, 16, and 24. The hidden states preceding the attention and MLP layers exhibit gaussian-like shapes, while the intermediate states within those blocks exhibit laplacian-like shapes. All distributions are zero-mean and unimodal, with the concentration around zero motivating magnitude-based pruning. The authors note that some activations are heavy-tailed and contain outliers, consistent with prior work.