SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jia, Zhao, Ziyu, Ni, Tingjuntao, Wei, Zhongyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915549615751168
author Wang, Jia
Zhao, Ziyu
Ni, Tingjuntao
Wei, Zhongyu
author_facet Wang, Jia
Zhao, Ziyu
Ni, Tingjuntao
Wei, Zhongyu
contents Large language models (LLMs) show strong potential for simulating human social behaviors and interactions, yet lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes. To bridge this gap, we introduce SocioBench-a comprehensive benchmark derived from the annually collected, standardized survey data of the International Social Survey Programme (ISSP). The benchmark aggregates over 480,000 real respondent records from more than 30 countries, spanning 10 sociological domains and over 40 demographic attributes. Our experiments indicate that LLMs achieve only 30-40% accuracy when simulating individuals in complex survey scenarios, with statistically significant differences across domains and demographic subgroups. These findings highlight several limitations of current LLMs in survey scenarios, including insufficient individual-level data coverage, inadequate scenario diversity, and missing group-level modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models
Wang, Jia
Zhao, Ziyu
Ni, Tingjuntao
Wei, Zhongyu
Social and Information Networks
Computers and Society
Large language models (LLMs) show strong potential for simulating human social behaviors and interactions, yet lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes. To bridge this gap, we introduce SocioBench-a comprehensive benchmark derived from the annually collected, standardized survey data of the International Social Survey Programme (ISSP). The benchmark aggregates over 480,000 real respondent records from more than 30 countries, spanning 10 sociological domains and over 40 demographic attributes. Our experiments indicate that LLMs achieve only 30-40% accuracy when simulating individuals in complex survey scenarios, with statistically significant differences across domains and demographic subgroups. These findings highlight several limitations of current LLMs in survey scenarios, including insufficient individual-level data coverage, inadequate scenario diversity, and missing group-level modeling.
title SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models
topic Social and Information Networks
Computers and Society
url https://arxiv.org/abs/2510.11131