IC-1552Deduplication of pretraining data does not consistently improve gender fairness in Pythia models

Catarina G Belém, Preethi Seshadri, Yasaman Razeghi, Sameer Singh

SourceAre Models Biased on Text without Gender-related Language?

The paper compares original and deduplicated Pythia models (70M-12B) on all five benchmarks using the US metric. The effect of deduplication is inconsistent across sizes and benchmarks: it improves fairness for Pythia-70M (+6.73 on USE-5) and Pythia-6.9B (+6.52 on USE-5) but worsens it for Pythia-410M (-17.14 on USE-5) and Pythia-12B (-11.65 on USE-5). The authors conclude that deduplication does not provide a reliable path to reducing gender bias in non-stereotypical settings.

Evidence
correlational
Key metric
USE-5 fairness change from deduplication: Pythia-70M +6.73, 160M -1.38, 410M -17.14, 1.4B -5.52, 2.8B +4.62, 6.9B +6.52, 12B -11.65; WG: 70M +2.80, 160M -3.74, 410M -1.87, 1.4B -18.69, 2.8B -1.87, 6.9B +9.35, 12B +1.87
Caveat
Only Pythia models have publicly available deduplicated variants, so the finding is limited to this one model family.
Model
Pythia
Datasets
USE-5 [eval], Winobias [eval], Winogender [eval]
Related findings
IC-1549, IC-1550, IC-1551
Extraction
automatic-extraction