Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-840
GPT-2 Small's name mover heads exhibit disrupted attention patterns under out-of-distribution Gaussian noise corruption