IC-36570B LLM variants tolerate substantially higher activation sparsity than smaller counterparts, and Llama-3 shows more degradation than Llama-2 and Mistral at 50% sparsity
James Liu, Pragaash Ponnusamy, Tianle Cai, Han Guo, Yoon Kim, Ben Athiwaratkun
Across Llama-2, Llama-3, and Mistral families at sizes 7B to 70B, the 70B models maintain performance at higher sparsity levels. At 50% model-wide sparsity, Llama-3-8B drops from 68.07 to 63.42 on the six-task average while Llama-2-70B drops only from 72.65 to 72.02. At 65% sparsity, most models degrade significantly, with the exception of Llama-2-70B which remains reasonably performant (69.30). The authors note this aligns with prior work showing quantization techniques are less effective on newer models trained on more tokens.