The paper compares original and deduplicated Pythia models (70M-12B) on all five benchmarks using the US metric. The effect of deduplication is inconsistent across sizes and benchmarks: it improves fairness for Pythia-70M (+6.73 on USE-5) and Pythia-6.9B (+6.52 on USE-5) but worsens it for Pythia-410M (-17.14 on USE-5) and Pythia-12B (-11.65 on USE-5). The authors conclude that deduplication does not provide a reliable path to reducing gender bias in non-stereotypical settings.