Across all 450 SOTOPIA tasks, every model tested (GPT-4, GPT-3.5, Llama-2-70B-Chat, MPT-30B-Chat) receives a negative average score on both the social rules (soc) and secret (sec) dimensions, indicating systematic failure to maintain social norms and keep private information. GPT-4, despite outperforming others on most dimensions, is not better than the other models on soc and sec. Qualitative examples show models leaking secrets at the beginning of conversations and violating social norms during interactions.
The scores are close to zero, indicating the violations are minor on average; the evaluation is automated by GPT-4, which the authors note tends to rate higher than humans on soc and sec dimensions.