IC-747SOTA pruning methods (SparseGPT, Wanda, magnitude) cause significant degradation on knowledge-intensive tasks for Vicuna and Llama models at 25-30%+ unstructured sparsity, and fail completely for n:m structured sparsity

AJAY KUMAR JAISWAL, Zhe Gan, Xianzhi Du, Bowen Zhang, Zhangyang Wang, Yinfei Yang

SourceCompressing LLMs: The Truth is Rarely Pure and Never Simple

The paper applies unstructured and structured (n:m) pruning to Vicuna-7B, Vicuna-13B, Llama-7B, and Llama-2-7B, then evaluates on FreebaseQA (factoid QA), MMLU (multiple-choice reasoning), and MT-Bench (instruction following). All pruning methods show negligible perplexity change up to 45-60% sparsity, yet suffer catastrophic drops on knowledge-intensive tasks: on FreebaseQA, magnitude-pruned Vicuna-7B drops from 65.44% to 13.99% at 50% sparsity, and on MMLU from 47.1% to 5.0%. For n:m structured sparsity, all methods show performance drops as severe as 50% or more. Quantization (GPTQ) is comparatively more successful, with 8-bit and 4-bit variants remaining matching for Vicuna-7B and Vicuna-13B respectively.

Evidence
interventional
Key metric
FreebaseQA exact match: magnitude 65.44→13.99, SparseGPT 65.44→42.86, Wanda 65.44→44.66 at 50% sparsity (Vicuna-7B); MMLU accuracy: magnitude 0.471→0.050, SparseGPT 0.471→0.308, Wanda 0.471→0.386 at 50% sparsity; n:m sparsity performance drop ≥50% for all methods
Caveat
Evaluation restricted to Vicuna (decoder-only) architecture due to open-source license; matching threshold set at 5% performance drop; results averaged across 3 independent runs
Model
Vicuna Vicuna-7B, Vicuna-13B, LLaMA Llama 7B, Llama 2 / Llama 2 base Llama 2 7B
Concepts
Failure mode
Datasets
FreebaseQA [eval], MMLU / MMLU-Math [eval], MT-Bench [eval]
Methods
SparseGPT [primary], Wanda [primary], Magnitude Pruning [primary], GPTQ [compared-to]
Related work
SparseGPT [builds-on], Wanda [builds-on], GPTQ [builds-on]
Related findings
IC-748, IC-749
Extraction
automatic-extraction