IC-252GPT models produce harmful gender stereotypes at higher rates when user names imply a demographic group, with GPT-3.5 Turbo showing the highest rates and open-ended generation tasks most affected

Tyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu, Anna-Luisa Brakman, Pamela Mishkin, Meghan Shah, Johannes Heidecke, Lilian Weng, Adam Tauman Kalai

SourceFirst-Person Fairness in Chatbots

The paper replays real user prompts with names from different gender groups and measures how often the resulting response pairs are rated as harmful stereotypes by an LMRA (GPT-4o). Across six GPT models and 66 tasks in 9 domains, harmful gender stereotype rates are all a fraction of 1% on average over domains, but the 'write a story' task exhibits the greatest rate of harms. GPT-3.5 Turbo shows the highest harm rate among all models. The LMRA ratings for gender bias strongly correlate with human crowd-worker ratings (ρ=0.86, sign agreement 90%).

Evidence
correlational
Key metric
harms on average over domains are all a fraction of 1%; write a story task exhibited the greatest rate of harms; gpt-3.5 exhibited the greatest harm rate; LMRA-human correlation ρ=0.86, a=90% for gender
Caveat
Results are not directly reproducible due to privacy; the LMRA is GPT-4o, a model from the same family as the chatbots being evaluated, which could introduce systematic agreement.
Model
GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, GPT-4o mini, O1 / OpenAI-o1-preview o1-preview, O1-mini
Concepts
Failure mode
Datasets
LMSYS-Chat-1M [source], WildChat [source]
Related work
Tamkin et al. 2023 [context], Nghiem et al. 2024 [context], Dwivedi-Yu et al. 2024 [context]
Related findings
IC-253, IC-254, IC-255
Extraction
automatic-extraction