IC-366In Llama-3-70B, activation sparsifiability varies systematically across depth: Wq/Wk peak in block 0 then decline sharply, Wo peaks at 80-90% mid-model, and Wdown is consistently more sparsifiable than Wgate and Wup
James Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo, Yoon Kim, Ben Athiwaratkun
Using block-wise greedy optimization at 50% model-level sparsity, the paper reveals depth-dependent structure in Llama-3-70B. Attention parameters Wq and Wk exhibit high sparsifiability in block 0 followed by a sharp decline. Wo's sparsifiability starts at 50-60%, peaks at 80-90% in mid-model blocks, then returns to 50-60% in the final blocks. For MLP parameters, Wdown is more sparsifiable than Wgate, which is more sparsifiable than Wup, across all blocks. The authors attribute Wdown's higher sparsifiability to its laplacian-shaped input distribution and Wgate's intermediate position to SiLU decreasing the saliency of negative outputs.
Evidence
observational
Key metric
Wo sparsifiability: 50-60% (early blocks) → 80-90% (mid-model) → 50-60% (final blocks); Wq/Wk: high in block 0 then sharp decline; Wdown > Wgate > Wup across all blocks
Caveat
The authors note these patterns 'tend to hold for the other models we analyze' but the detailed plots are for Llama-3-70B specifically.