IC-1388In LLaMA-13B, feed-forward layers are more critical than attention layers for fact recall, while both are equally important for in-context learning
Tian Jin, Nolan Clement, Xin Dong, Vaishnavh Nagarajan, Michael Carbin, Jonathan Ragan-Kelley, Gintare Karolina Dziugaite
The authors selectively prune only attention layers or only feed-forward (FFW) layers in LLaMA-13B and measure the effect on fact recall (close-book TriviaQA) and ICL (linear classification). Pruning 60% of FFW parameters causes 14% more accuracy degradation on fact recall than pruning 60% of attention parameters. For ICL, the two module types show similar importance. This suggests fact recall is stored more heavily in FFW layers while ICL relies on both.
Evidence
interventional
Key metric
"pruning 60% ffw layers lead to 14% more accuracy degradation than pruning 60% of attention layers" on TriviaQA close-book
Caveat
"preliminary results" (as stated in the main text); single model (LLaMA-13B only)