IC-1404On SOTOPIA-HARD, GPT-4 achieves significantly lower goal completion than humans and exhibits non-strategic negotiation and excessive compromise behaviors

Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, Maarten Sap

SourceSOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

On the 20 most challenging SOTOPIA tasks (SOTOPIA-HARD), GPT-4 scores 4.85 on goal completion versus 5.95 for humans interacting with GPT-4 and 6.15 for human-human interactions (p<0.05). GPT-4 produces 45.5 words per turn versus 16.8 for humans, indicating less efficient communication. Qualitatively, GPT-4 always rephrases the other agent's utterance before answering (active listening from RLHF training), starts bargaining at its exact target price rather than strategically, and tends to propose compromised solutions rather than persisting in its stated goals.

Evidence
correlational
Key metric
Goal: GPT-4 4.85 vs human(w GPT-4) 5.95* vs human(w human) 6.15*; words/turn: GPT-4 45.5 vs human 16.8; *p<0.05
Caveat
Only 20 tasks in SOTOPIA-HARD; 40 human-GPT-4 and 20 human-human interactions; evaluation by both GPT-4 and human annotators shows similar trends.
Model
GPT-4 / ChatGPT4 / GPT-4 Code Interpreter / GPT-4 Technical Report
Concepts
Failure mode
Datasets
SOTOPIA [eval]
Methods
SOTOPIA-EVAL [primary]
Related findings
IC-1401, IC-1402, IC-1403
Extraction
automatic-extraction