IC-253Post-training reinforcement learning significantly reduces harmful gender stereotypes in GPT models, with the best-fit slope of 0.21 indicating post-RL models have far lower bias than pre-RL versions

Tyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu, Anna-Luisa Brakman, Pamela Mishkin, Meghan Shah, Johannes Heidecke, Lilian Weng, Adam Tauman Kalai

SourceFirst-Person Fairness in Chatbots

The paper compares harmful gender stereotype ratings for four GPT models before and after post-training RL across 19 selected tasks. All post-RL models fall below the 45-degree line, indicating reduced bias. The best-fit slope is 0.21 (95% CI: 0.17, 0.24), meaning post-RL models retain only about 21% of the pre-RL bias level. Individual slopes range from 0.08 (GPT-4o Mini) to 0.37 (GPT-4 Turbo), showing the effect is consistent across the model family.

Evidence
correlational
Key metric
slope of best-fit line is 0.21 (95% ci: 0.17, 0.24); individual slopes: gpt-3.5t (slope=0.31), gpt-4t (slope=0.37), gpt-4o (slope=0.26), gpt-4o-mini (slope=0.08)
Caveat
The authors note it is 'rl (and possibly other post-sft mitigations)' that reduce bias, so the effect is not isolated to RL alone. Pre-RL models are intermediate checkpoints not independently released.
Model
GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4 Turbo, GPT-4o mini
Datasets
LMSYS-Chat-1M [source], WildChat [source]
Related work
Achiam et al. 2023 [context], Perez et al. 2023 [context]
Related findings
IC-252, IC-254, IC-255
Extraction
automatic-extraction