IC-366In Llama-3-70B, activation sparsifiability varies systematically across depth: Wq/Wk peak in block 0 then decline sharply, Wo peaks at 80-90% mid-model, and Wdown is consistently more sparsifiable than Wgate and Wup

James Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo, Yoon Kim, Ben Athiwaratkun

SourceTraining-Free Activation Sparsity in Large Language Models

Using block-wise greedy optimization at 50% model-level sparsity, the paper reveals depth-dependent structure in Llama-3-70B. Attention parameters Wq and Wk exhibit high sparsifiability in block 0 followed by a sharp decline. Wo's sparsifiability starts at 50-60%, peaks at 80-90% in mid-model blocks, then returns to 50-60% in the final blocks. For MLP parameters, Wdown is more sparsifiable than Wgate, which is more sparsifiable than Wup, across all blocks. The authors attribute Wdown's higher sparsifiability to its laplacian-shaped input distribution and Wgate's intermediate position to SiLU decreasing the saliency of negative outputs.

Evidence
observational
Key metric
Wo sparsifiability: 50-60% (early blocks) → 80-90% (mid-model) → 50-60% (final blocks); Wq/Wk: high in block 0 then sharp decline; Wdown > Wgate > Wup across all blocks
Caveat
The authors note these patterns 'tend to hold for the other models we analyze' but the detailed plots are for Llama-3-70B specifically.
Model
Llama 3 70B
Concepts
Depth-dependent structure
Datasets
C4 [source]
Related work
FinerCut [context]
Related findings
IC-364, IC-365, IC-367
Extraction
automatic-extraction