IC-36570B LLM variants tolerate substantially higher activation sparsity than smaller counterparts, and Llama-3 shows more degradation than Llama-2 and Mistral at 50% sparsity

James Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo, Yoon Kim, Ben Athiwaratkun

SourceTraining-Free Activation Sparsity in Large Language Models

Across Llama-2, Llama-3, and Mistral families at sizes 7B to 70B, the 70B models maintain performance at higher sparsity levels. At 50% model-wide sparsity, Llama-3-8B drops from 68.07 to 63.42 on the six-task average while Llama-2-70B drops only from 72.65 to 72.02. At 65% sparsity, most models degrade significantly, with the exception of Llama-2-70B which remains reasonably performant (69.30). The authors note this aligns with prior work showing quantization techniques are less effective on newer models trained on more tokens.

Evidence
correlational
Key metric
Llama-3-8b downstream avg: 68.07 (0%) → 63.42 (50%) → 52.59 (65%); Llama-2-70b: 72.65 → 72.02 → 69.30; Llama-3-70b: 80.41 → 78.26 → 73.07; Llama-2-7b: 56.50 → 54.26 → 48.16; Mistral-7b: 66.96 → 64.16 → 58.93
Caveat
Results between Llama-3 and Llama-2/Mistral are not directly comparable due to differing vocabulary sizes.
Model
Llama 3 8B, 70B, Llama 2 / Llama 2 base Llama 2 7B, Llama 2 13B, Llama 2 70B, Mistral 7B / Mistral / Mistral 3 7B / Mistral-0.2-7B / Mistral-v0.1
Concepts
Scale-dependent behaviour
Datasets
WikiText [eval], MMLU / MMLU-Math [eval], ARC-Challenge [eval], HellaSwag [eval], GSM8K [eval], PIQA [eval], Winogrande [eval]
Methods
CATS [compared-to]
Related work
CATS [compared-to]
Related findings
IC-364, IC-366, IC-367
Extraction
automatic-extraction