IC-252GPT models produce harmful gender stereotypes at higher rates when user names imply a demographic group, with GPT-3.5 Turbo showing the highest rates and open-ended generation tasks most affected
Tyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu, Anna-Luisa Brakman, Pamela Mishkin, Meghan Shah, Johannes Heidecke, Lilian Weng, Adam Tauman Kalai
The paper replays real user prompts with names from different gender groups and measures how often the resulting response pairs are rated as harmful stereotypes by an LMRA (GPT-4o). Across six GPT models and 66 tasks in 9 domains, harmful gender stereotype rates are all a fraction of 1% on average over domains, but the 'write a story' task exhibits the greatest rate of harms. GPT-3.5 Turbo shows the highest harm rate among all models. The LMRA ratings for gender bias strongly correlate with human crowd-worker ratings (ρ=0.86, sign agreement 90%).
Evidence
correlational
Key metric
harms on average over domains are all a fraction of 1%; write a story task exhibited the greatest rate of harms; gpt-3.5 exhibited the greatest harm rate; LMRA-human correlation ρ=0.86, a=90% for gender
Caveat
Results are not directly reproducible due to privacy; the LMRA is GPT-4o, a model from the same family as the chatbots being evaluated, which could introduce systematic agreement.