IC-254GPT-4o Mini responses to female-sounding names systematically use simpler, more light-hearted, and less technical language compared to male-sounding names across multiple task domains

Tyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu, Anna-Luisa Brakman, Pamela Mishkin, Meghan Shah, Johannes Heidecke, Lilian Weng, Adam Tauman Kalai

SourceFirst-Person Fairness in Chatbots

Using the bias enumeration algorithm, the paper identifies axes of difference in GPT-4o Mini responses across 100k prompts. For female-sounding names, 52-55% of prompts elicit responses rated as simpler, more light-hearted, or avoiding technical terms. The most significant group-a (female) biased axes are: uses more general and layman-friendly language (53%), gives simpler explanations (53%), and generally gives concise straightforward responses (52%). In the 'write a story' task, female names prompted stories with female main characters more often. The LMRA ratings of these features were only weakly correlated with human ratings.

Evidence
correlational
Key metric
52-55% of prompts result in responses for f-names rated as simpler, more light-hearted, or avoid technical terms; uses more general and layman-friendly language (53%); gives simpler explanations (53%); generally gives concise, straightforward responses and explanations (52%)
Caveat
The LMRA ratings of features such as simple language were only weakly correlated with human ratings, reducing confidence in the specific axes identified.
Model
GPT-4o mini
Concepts
Failure mode
Datasets
LMSYS-Chat-1M [source], WildChat [source]
Related work
Zhong et al. 2022 [builds-on], Findeis et al. 2024 [builds-on]
Related findings
IC-252, IC-253, IC-255
Extraction
automatic-extraction