SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Zixiang, Wang, Yanbo, Huang, Yue, Ye, Jiayi, Zhuang, Haomin, Song, Zirui, Gao, Lang, Wang, Chenxi, Chen, Zhaorun, Zhou, Yujun, Li, Sixian, Pan, Wang, Zhao, Yue, Zhao, Jieyu, Zhang, Xiangliang, Chen, Xiuying
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913866154246144
author Xu, Zixiang
Wang, Yanbo
Huang, Yue
Ye, Jiayi
Zhuang, Haomin
Song, Zirui
Gao, Lang
Wang, Chenxi
Chen, Zhaorun
Zhou, Yujun
Li, Sixian
Pan, Wang
Zhao, Yue
Zhao, Jieyu
Zhang, Xiangliang
Chen, Xiuying
author_facet Xu, Zixiang
Wang, Yanbo
Huang, Yue
Ye, Jiayi
Zhuang, Haomin
Song, Zirui
Gao, Lang
Wang, Chenxi
Chen, Zhaorun
Zhou, Yujun
Li, Sixian
Pan, Wang
Zhao, Yue
Zhao, Jieyu
Zhang, Xiangliang
Chen, Xiuying
contents Large language models (LLMs) are increasingly applied to socially grounded tasks, such as online community moderation, media content analysis, and social reasoning games. Success in these contexts depends on a model's social reasoning ability - the capacity to interpret social contexts, infer others' mental states, and assess the truthfulness of presented information. However, there is currently no systematic evaluation framework that comprehensively assesses the social reasoning capabilities of LLMs. Existing efforts often oversimplify real-world scenarios and consist of tasks that are too basic to challenge advanced models. To address this gap, we introduce SocialMaze, a new benchmark specifically designed to evaluate social reasoning. SocialMaze systematically incorporates three core challenges: deep reasoning, dynamic interaction, and information uncertainty. It provides six diverse tasks across three key settings: social reasoning games, daily-life interactions, and digital community platforms. Both automated and human validation are used to ensure data quality. Our evaluation reveals several key insights: models vary substantially in their ability to handle dynamic interactions and integrate temporally evolving information; models with strong chain-of-thought reasoning perform better on tasks requiring deeper inference beyond surface-level cues; and model reasoning degrades significantly under uncertainty. Furthermore, we show that targeted fine-tuning on curated reasoning examples can greatly improve model performance in complex social scenarios. The dataset is publicly available at: https://huggingface.co/datasets/MBZUAI/SocialMaze
format Preprint
id arxiv_https___arxiv_org_abs_2505_23713
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
Xu, Zixiang
Wang, Yanbo
Huang, Yue
Ye, Jiayi
Zhuang, Haomin
Song, Zirui
Gao, Lang
Wang, Chenxi
Chen, Zhaorun
Zhou, Yujun
Li, Sixian
Pan, Wang
Zhao, Yue
Zhao, Jieyu
Zhang, Xiangliang
Chen, Xiuying
Computation and Language
Large language models (LLMs) are increasingly applied to socially grounded tasks, such as online community moderation, media content analysis, and social reasoning games. Success in these contexts depends on a model's social reasoning ability - the capacity to interpret social contexts, infer others' mental states, and assess the truthfulness of presented information. However, there is currently no systematic evaluation framework that comprehensively assesses the social reasoning capabilities of LLMs. Existing efforts often oversimplify real-world scenarios and consist of tasks that are too basic to challenge advanced models. To address this gap, we introduce SocialMaze, a new benchmark specifically designed to evaluate social reasoning. SocialMaze systematically incorporates three core challenges: deep reasoning, dynamic interaction, and information uncertainty. It provides six diverse tasks across three key settings: social reasoning games, daily-life interactions, and digital community platforms. Both automated and human validation are used to ensure data quality. Our evaluation reveals several key insights: models vary substantially in their ability to handle dynamic interactions and integrate temporally evolving information; models with strong chain-of-thought reasoning perform better on tasks requiring deeper inference beyond surface-level cues; and model reasoning degrades significantly under uncertainty. Furthermore, we show that targeted fine-tuning on curated reasoning examples can greatly improve model performance in complex social scenarios. The dataset is publicly available at: https://huggingface.co/datasets/MBZUAI/SocialMaze
title SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2505.23713