On the 20 most challenging SOTOPIA tasks (SOTOPIA-HARD), GPT-4 scores 4.85 on goal completion versus 5.95 for humans interacting with GPT-4 and 6.15 for human-human interactions (p<0.05). GPT-4 produces 45.5 words per turn versus 16.8 for humans, indicating less efficient communication. Qualitatively, GPT-4 always rephrases the other agent's utterance before answering (active listening from RLHF training), starts bargaining at its exact target price rather than strategically, and tends to propose compromised solutions rather than persisting in its stated goals.
Evidence
correlational
Key metric
Goal: GPT-4 4.85 vs human(w GPT-4) 5.95* vs human(w human) 6.15*; words/turn: GPT-4 45.5 vs human 16.8; *p<0.05
Caveat
Only 20 tasks in SOTOPIA-HARD; 40 human-GPT-4 and 20 human-human interactions; evaluation by both GPT-4 and human annotators shows similar trends.