Modelpedia
Work in progress
About
Findings
Models
Concepts
Methods
Datasets
Sources
Light
Dark
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
2024-01-16
· ICLR 2024 spotlight ·
anchor
Findings
IC-1401
GPT-4 serves as a proxy for human judgment on SOTOPIA-EVAL, with strong correlations on goal, financial, and relationship dimensions for model outputs
IC-1402
Llama-2-70B-Chat underperforms GPT-3.5 across all SOTOPIA dimensions in interactive social scenarios, diverging from static benchmark rankings
IC-1403
All four evaluated LLMs produce negative scores on social rules and secret-keeping dimensions in SOTOPIA interactions
IC-1404
On SOTOPIA-HARD, GPT-4 achieves significantly lower goal completion than humans and exhibits non-strategic negotiation and excessive compromise behaviors