AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mou, Xinyi, Liang, Jingcong, Lin, Jiayu, Zhang, Xinnong, Liu, Xiawei, Yang, Shiyue, Ye, Rong, Chen, Lei, Kuang, Haoyu, Huang, Xuanjing, Wei, Zhongyu
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909400681152512
author Mou, Xinyi
Liang, Jingcong
Lin, Jiayu
Zhang, Xinnong
Liu, Xiawei
Yang, Shiyue
Ye, Rong
Chen, Lei
Kuang, Haoyu
Huang, Xuanjing
Wei, Zhongyu
author_facet Mou, Xinyi
Liang, Jingcong
Lin, Jiayu
Zhang, Xinnong
Liu, Xiawei
Yang, Shiyue
Ye, Rong
Chen, Lei
Kuang, Haoyu
Huang, Xuanjing
Wei, Zhongyu
contents Large language models (LLMs) are increasingly leveraged to empower autonomous agents to simulate human beings in various fields of behavioral research. However, evaluating their capacity to navigate complex social interactions remains a challenge. Previous studies face limitations due to insufficient scenario diversity, complexity, and a single-perspective focus. To this end, we introduce AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios. Drawing on Dramaturgical Theory, AgentSense employs a bottom-up approach to create 1,225 diverse social scenarios constructed from extensive scripts. We evaluate LLM-driven agents through multi-turn interactions, emphasizing both goal completion and implicit reasoning. We analyze goals using ERG theory and conduct comprehensive experiments. Our findings highlight that LLMs struggle with goals in complex social scenarios, especially high-level growth needs, and even GPT-4o requires improvement in private information reasoning. Code and data are available at \url{https://github.com/ljcleo/agent_sense}.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19346
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
Mou, Xinyi
Liang, Jingcong
Lin, Jiayu
Zhang, Xinnong
Liu, Xiawei
Yang, Shiyue
Ye, Rong
Chen, Lei
Kuang, Haoyu
Huang, Xuanjing
Wei, Zhongyu
Computation and Language
Computers and Society
Large language models (LLMs) are increasingly leveraged to empower autonomous agents to simulate human beings in various fields of behavioral research. However, evaluating their capacity to navigate complex social interactions remains a challenge. Previous studies face limitations due to insufficient scenario diversity, complexity, and a single-perspective focus. To this end, we introduce AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios. Drawing on Dramaturgical Theory, AgentSense employs a bottom-up approach to create 1,225 diverse social scenarios constructed from extensive scripts. We evaluate LLM-driven agents through multi-turn interactions, emphasizing both goal completion and implicit reasoning. We analyze goals using ERG theory and conduct comprehensive experiments. Our findings highlight that LLMs struggle with goals in complex social scenarios, especially high-level growth needs, and even GPT-4o requires improvement in private information reasoning. Code and data are available at \url{https://github.com/ljcleo/agent_sense}.
title AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2410.19346