SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Xuhui, Zhu, Hao, Mathur, Leena, Zhang, Ruohong, Yu, Haofei, Qi, Zhengyang, Morency, Louis-Philippe, Bisk, Yonatan, Fried, Daniel, Neubig, Graham, Sap, Maarten
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916172779225088
author Zhou, Xuhui
Zhu, Hao
Mathur, Leena
Zhang, Ruohong
Yu, Haofei
Qi, Zhengyang
Morency, Louis-Philippe
Bisk, Yonatan
Fried, Daniel
Neubig, Graham
Sap, Maarten
author_facet Zhou, Xuhui
Zhu, Hao
Mathur, Leena
Zhang, Ruohong
Yu, Haofei
Qi, Zhengyang
Morency, Louis-Philippe
Bisk, Yonatan
Fried, Daniel
Neubig, Graham
Sap, Maarten
contents Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.
format Preprint
id arxiv_https___arxiv_org_abs_2310_11667
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Zhou, Xuhui
Zhu, Hao
Mathur, Leena
Zhang, Ruohong
Yu, Haofei
Qi, Zhengyang
Morency, Louis-Philippe
Bisk, Yonatan
Fried, Daniel
Neubig, Graham
Sap, Maarten
Artificial Intelligence
Computation and Language
Machine Learning
Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents.
title SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2310.11667