IC-1403All four evaluated LLMs produce negative scores on social rules and secret-keeping dimensions in SOTOPIA interactions

Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, Maarten Sap

SourceSOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

Across all 450 SOTOPIA tasks, every model tested (GPT-4, GPT-3.5, Llama-2-70B-Chat, MPT-30B-Chat) receives a negative average score on both the social rules (soc) and secret (sec) dimensions, indicating systematic failure to maintain social norms and keep private information. GPT-4, despite outperforming others on most dimensions, is not better than the other models on soc and sec. Qualitative examples show models leaking secrets at the beginning of conversations and violating social norms during interactions.

Evidence
correlational
Key metric
Soc: GPT-4 -0.07, GPT-3.5 -0.08, Llama-2 -0.11, MPT -0.09; Sec: GPT-4 -0.14, GPT-3.5 -0.08, Llama-2 -0.14, MPT -0.07
Caveat
The scores are close to zero, indicating the violations are minor on average; the evaluation is automated by GPT-4, which the authors note tends to rate higher than humans on soc and sec dimensions.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report, GPT-3.5 / ChatGPT-3.5, Llama 2 / Llama 2 base Llama-2-70B-Chat, MPT MPT-30B-Chat
Concepts
Failure mode
Datasets
SOTOPIA [eval]
Methods
SOTOPIA-EVAL [primary]
Related findings
IC-1401, IC-1402, IC-1404
Extraction
automatic-extraction