The paper administered the 57-item PVQ-RR questionnaire to six LLMs under five prompting strategies and two temperatures, then computed Spearman rank correlations between each model's value ranking and the human benchmark from Schwartz and Cieciuch (2022, 49 cultural groups, N=53,472). Most prompt-model combinations yielded correlations above 0.8 with p < .001. The value anchor prompt produced correlations ranging from 0.75 to 0.85. A notable exception was GPT-4-0314 under the basic prompt, which showed very low correlation and stark deviations (conformity-rules ranked first vs. humans' 16th; benevolence-care ranked 18th vs. humans' 1st).
Evidence
correlational
Key metric
Spearman rank correlations > 0.8 for many prompt-model combinations (p < .001); value anchor prompt correlations 0.75 to 0.85
Caveat
GPT-4-0314 under the basic prompt shows very low correlation with the human hierarchy, deviating sharply (conformity-rules ranked 1st vs. humans' 16th; benevolence-care ranked 18th vs. humans' 1st).