IC-255GPT-4o, Llama 3.1, and Claude models show varying correlation with human ratings when used as stereotype evaluators, with GPT-4o achieving the strongest gender correlation (ρ=0.86) but weaker racial correlations

Tyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu, Anna-Luisa Brakman, Pamela Mishkin, Meghan Shah, Johannes Heidecke, Lilian Weng, Adam Tauman Kalai

SourceFirst-Person Fairness in Chatbots

The paper evaluates how well different LMs, when used as LMRAs, agree with human crowd-worker ratings of harmful stereotypes. For gender bias, GPT-4o achieves ρ=0.86 (sign agreement 90%), Llama 3.1 70B achieves ρ=0.84 (88%), and Claude 3.5 Sonnet achieves ρ=0.85 (88%). For racial bias, correlations are substantially weaker across all models, with GPT-4o at ρ=0.75 (Asian-White) down to ρ=0.34 (Hispanic-White). Smaller models like Llama 3.1 8B show much weaker agreement (ρ=0.26 for gender). LMAs from other families do not show substantially higher agreement than GPT-4o.

Evidence
correlational
Key metric
4o (ours) ρ=0.86 a=90% (gender); l3.1 8b ρ=0.26 a=52%; l3.1 70b ρ=0.84 a=88%; l3.1 405b ρ=0.82 a=87%; c3.5 haiku ρ=0.72 a=58%; c3.5 sonnet ρ=0.85 a=88%; c3 opus ρ=0.62 a=29%
Caveat
The LMRA is run at temperature 0 while the chatbot responses are at temperature 0.8; the zero-shot LMRA instructions may underperform compared to few-shot or calibrated approaches; racial bias correlations are notably weaker than gender.
Model
GPT-4o, Llama 3.1 8B, 70B, 405B, Claude 3.5 Haiku, Sonnet, Claude 3 Opus
Datasets
LMSYS-Chat-1M [source], WildChat [source]
Related work
Perez et al. 2023 [context], Liu et al. 2024 [context]
Related findings
IC-252, IC-253, IC-254
Extraction
automatic-extraction