Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
First-Person Fairness in Chatbots
2025-01-22
· ICLR 2025 Spotlight ·
anchor
Findings
IC-252
GPT models produce harmful gender stereotypes at higher rates when user names imply a demographic group, with GPT-3.5 Turbo showing the highest rates and open-ended generation tasks most affected
IC-253
Post-training reinforcement learning significantly reduces harmful gender stereotypes in GPT models, with the best-fit slope of 0.21 indicating post-RL models have far lower bias than pre-RL versions
IC-254
GPT-4o Mini responses to female-sounding names systematically use simpler, more light-hearted, and less technical language compared to male-sounding names across multiple task domains
IC-255
GPT-4o, Llama 3.1, and Claude models show varying correlation with human ratings when used as stereotype evaluators, with GPT-4o achieving the strongest gender correlation (ρ=0.86) but weaker racial correlations