Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language Models
2024-01-16
· ICLR 2024 poster ·
anchor
Findings
IC-1005
Social bias neurons in BERT-base-cased and RoBERTa-base are concentrated in the deepest transformer layers