IC-255GPT-4o, Llama 3.1, and Claude models show varying correlation with human ratings when used as stereotype evaluators, with GPT-4o achieving the strongest gender correlation (ρ=0.86) but weaker racial correlations
Tyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu, Anna-Luisa Brakman, Pamela Mishkin, Meghan Shah, Johannes Heidecke, Lilian Weng, Adam Tauman Kalai
The paper evaluates how well different LMs, when used as LMRAs, agree with human crowd-worker ratings of harmful stereotypes. For gender bias, GPT-4o achieves ρ=0.86 (sign agreement 90%), Llama 3.1 70B achieves ρ=0.84 (88%), and Claude 3.5 Sonnet achieves ρ=0.85 (88%). For racial bias, correlations are substantially weaker across all models, with GPT-4o at ρ=0.75 (Asian-White) down to ρ=0.34 (Hispanic-White). Smaller models like Llama 3.1 8B show much weaker agreement (ρ=0.26 for gender). LMAs from other families do not show substantially higher agreement than GPT-4o.
The LMRA is run at temperature 0 while the chatbot responses are at temperature 0.8; the zero-shot LMRA instructions may underperform compared to few-shot or calibrated approaches; racial bias correlations are notably weaker than gender.