IC-431Llama-2-7b generates more toxic content for female-associated prompts than male-associated prompts on the BOLD dataset

Ching-Yun Ko, Pin-Yu Chen, Payel Das, Youssef Mroueh, Soham Dan, Georgios Kollias, Subhajit Chaudhury, Tejaswini Pedapati, Luca Daniel

SourceLarge Language Models can Become Strong Self-Detoxifiers

Using the BOLD gender domain (776 'American actresses' vs 776 'American actors' prompts), the paper measures the toxic rate of Llama-2-7b's raw generations. The female-associated prompts yield a toxic rate of 0.066 and avg. max toxicity of 0.243, compared to 0.031 and 0.213 for male-associated prompts. The authors interpret this as the model being 'somewhat biased against female.' Both RAD and SASA mitigate the gap, with SASA reaching 0.027 vs 0.023 avg. max toxicity for female vs male.

Evidence
correlational
Key metric
toxic rate 0.066 (female) vs 0.031 (male); avg. max toxicity 0.243 (female) vs 0.213 (male) for raw Llama-2-7b on BOLD gender domain
Caveat
The authors note this is a 'somewhat' biased result and the gap is modest; the finding is reported in the appendix (Table 13) as a secondary analysis.
Model
Llama 2 / Llama 2 base Llama 2 7B
Concepts
Failure mode
Datasets
BOLD [eval]
Methods
Perspective API [eval]
Related findings
IC-430
Extraction
automatic-extraction