IC-827LLMs exhibit distinct psychological profiles that differ from human norms and vary by model size and version

Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho LAM, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, Michael Lyu

SourceOn the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs

Using thirteen psychometric scales across four domains (personality traits, interpersonal relationships, motivational tests, emotional abilities), the paper measures five released LLMs and compares them to human norms. LLMs generally score higher than humans on openness, conscientiousness, extraversion, emotional intelligence, and self-efficacy, while scoring lower on implicit culture belief. Model size matters: LLaMA-2-7B scores higher on GSE (39.1) than LLaMA-2-13B (30.4), and GPT-4 scores higher on GSE (39.9) than GPT-3.5-turbo (38.5). Most LLMs except text-davinci-003 and GPT-4 show elevated dark triad scores, and all show high lying subscale scores.

Evidence
correlational
Key metric
GSE: LLaMA-2-7B 39.1±1.2, LLaMA-2-13B 30.4±3.6, GPT-4 39.9±0.3, human 29.6±5.3; EIS: GPT-4 151.4±18.7, human 124.8±16.5; ICB: GPT-4 1.9±0.4, human 3.7±0.8; BFI openness: LLaMA-2-7B 4.2±0.3, human 3.9±0.7
Caveat
Human comparison data come from different demographic groups across different studies and countries, not a single representative global sample. The framework uses only Likert scales, excluding other psychometric methods.
Model
GPT-3 / GPT base text-davinci-003, GPT-3.5 / ChatGPT-3.5 GPT-3.5-turbo, GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, Llama 2 / Llama 2 base Llama 2 7B, Llama 2 13B
Concepts
Scale-dependent behaviour
Datasets
TruthfulQA / TruthfulQA MC1 [eval], SafetyQA [eval]
Related work
Miotto et al. 2022 [context], Wang et al. 2023a [context], Huang et al. 2023b [context]
Related findings
IC-828, IC-829
Extraction
automatic-extraction