Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
The Cost of Scaling Down Large Language Models: Reducing Model Size Affects Memory before In-context Learning
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1386
Fact recall in OPT and LLaMA models degrades by more than 5% relative accuracy when more than 30% of weights are pruned, and similarly when moving from the 30B to the 13B dense model
IC-1387
In-context learning capabilities in OPT and LLaMA models remain within 5% of dense-model accuracy even at 60-70% sparsity, and show less than 2% difference between the 30B and 1.3B dense OPT models
IC-1388
In LLaMA-13B, feed-forward layers are more critical than attention layers for fact recall, while both are equally important for in-context learning