SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916172779225088 |
|---|---|
| author | Zhou, Xuhui Zhu, Hao Mathur, Leena Zhang, Ruohong Yu, Haofei Qi, Zhengyang Morency, Louis-Philippe Bisk, Yonatan Fried, Daniel Neubig, Graham Sap, Maarten |
| author_facet | Zhou, Xuhui Zhu, Hao Mathur, Leena Zhang, Ruohong Yu, Haofei Qi, Zhengyang Morency, Louis-Philippe Bisk, Yonatan Fried, Daniel Neubig, Graham Sap, Maarten |
| contents | Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_11667 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents Zhou, Xuhui Zhu, Hao Mathur, Leena Zhang, Ruohong Yu, Haofei Qi, Zhengyang Morency, Louis-Philippe Bisk, Yonatan Fried, Daniel Neubig, Graham Sap, Maarten Artificial Intelligence Computation and Language Machine Learning Humans are social beings; we pursue social goals in our daily interactions, which is a crucial aspect of social intelligence. Yet, AI systems' abilities in this realm remain elusive. We present SOTOPIA, an open-ended environment to simulate complex social interactions between artificial agents and evaluate their social intelligence. In our environment, agents role-play and interact under a wide variety of scenarios; they coordinate, collaborate, exchange, and compete with each other to achieve complex social goals. We simulate the role-play interaction between LLM-based agents and humans within this task space and evaluate their performance with a holistic evaluation framework called SOTOPIA-Eval. With SOTOPIA, we find significant differences between these models in terms of their social intelligence, and we identify a subset of SOTOPIA scenarios, SOTOPIA-hard, that is generally challenging for all models. We find that on this subset, GPT-4 achieves a significantly lower goal completion rate than humans and struggles to exhibit social commonsense reasoning and strategic communication skills. These findings demonstrate SOTOPIA's promise as a general platform for research on evaluating and improving social intelligence in artificial agents. |
| title | SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents |
| topic | Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2310.11667 |