Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
USE-5
Findings
IC-1549
All 28 evaluated LMs exhibit gender bias on non-stereotypical sentence pairs, with fairness scores between 9% and 41%
[eval]
IC-1550
All evaluated LMs systematically prefer male pronoun completions in the non-stereotypical portions of Winobias and Winogender, with margins exceeding 40%
[eval]
IC-1551
No consistent relationship between model size and gender fairness scores is observed across six LM families
[eval]
IC-1552
Deduplication of pretraining data does not consistently improve gender fairness in Pythia models
[eval]