IC-103LLMs' value rankings align with the universal human value hierarchy under most prompting conditions

Naama Rozen, Liat Bezalel, Gal Elidan, Amir Globerson, Ella Daniel

SourceDo LLMs have Consistent Values?

The paper administered the 57-item PVQ-RR questionnaire to six LLMs under five prompting strategies and two temperatures, then computed Spearman rank correlations between each model's value ranking and the human benchmark from Schwartz and Cieciuch (2022, 49 cultural groups, N=53,472). Most prompt-model combinations yielded correlations above 0.8 with p < .001. The value anchor prompt produced correlations ranging from 0.75 to 0.85. A notable exception was GPT-4-0314 under the basic prompt, which showed very low correlation and stark deviations (conformity-rules ranked first vs. humans' 16th; benevolence-care ranked 18th vs. humans' 1st).

Evidence
correlational
Key metric
Spearman rank correlations > 0.8 for many prompt-model combinations (p < .001); value anchor prompt correlations 0.75 to 0.85
Caveat
GPT-4-0314 under the basic prompt shows very low correlation with the human hierarchy, deviating sharply (conformity-rules ranked 1st vs. humans' 16th; benevolence-care ranked 18th vs. humans' 1st).
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report GPT-4-0314, Gemini 1.0 Pro, Llama 3.1 8B, 70B, Gemma 2 9B, 27B
Datasets
PVQ-RR [eval]
Methods
PVQ-RR [eval], Spearman rank correlation [eval]
Related findings
IC-104, IC-105
Extraction
automatic-extraction