IC-1388In LLaMA-13B, feed-forward layers are more critical than attention layers for fact recall, while both are equally important for in-context learning

Tian Jin, Nolan Clement, Xin Dong, Vaishnavh Nagarajan, Michael Carbin, Jonathan Ragan-Kelley, Gintare Karolina Dziugaite

SourceThe Cost of Scaling Down Large Language Models: Reducing Model Size Affects Memory before In-context Learning

The authors selectively prune only attention layers or only feed-forward (FFW) layers in LLaMA-13B and measure the effect on fact recall (close-book TriviaQA) and ICL (linear classification). Pruning 60% of FFW parameters causes 14% more accuracy degradation on fact recall than pruning 60% of attention parameters. For ICL, the two module types show similar importance. This suggests fact recall is stored more heavily in FFW layers while ICL relies on both.

Evidence
interventional
Key metric
"pruning 60% ffw layers lead to 14% more accuracy degradation than pruning 60% of attention layers" on TriviaQA close-book
Caveat
"preliminary results" (as stated in the main text); single model (LLaMA-13B only)
Model
LLaMA
Datasets
TriviaQA [eval]
Methods
SparseGPT [primary]
Related work
Dai et al. 2021 (Knowledge neurons in pretrained transformers) [context]
Related findings
IC-1386, IC-1387
Extraction
automatic-extraction