Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yue, Yuan, Zhengqing, Zhou, Yujun, Guo, Kehan, Wang, Xiangqi, Zhuang, Haomin, Sun, Weixiang, Sun, Lichao, Wang, Jindong, Ye, Yanfang, Zhang, Xiangliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913568411090944
author Huang, Yue
Yuan, Zhengqing
Zhou, Yujun
Guo, Kehan
Wang, Xiangqi
Zhuang, Haomin
Sun, Weixiang
Sun, Lichao
Wang, Jindong
Ye, Yanfang
Zhang, Xiangliang
author_facet Huang, Yue
Yuan, Zhengqing
Zhou, Yujun
Guo, Kehan
Wang, Xiangqi
Zhuang, Haomin
Sun, Weixiang
Sun, Lichao
Wang, Jindong
Ye, Yanfang
Zhang, Xiangliang
contents Large Language Models (LLMs) are increasingly employed for simulations, enabling applications in role-playing agents and Computational Social Science (CSS). However, the reliability of these simulations is under-explored, which raises concerns about the trustworthiness of LLMs in these applications. In this paper, we aim to answer ``How reliable is LLM-based simulation?'' To address this, we introduce TrustSim, an evaluation dataset covering 10 CSS-related topics, to systematically investigate the reliability of the LLM simulation. We conducted experiments on 14 LLMs and found that inconsistencies persist in the LLM-based simulated roles. In addition, the consistency level of LLMs does not strongly correlate with their general performance. To enhance the reliability of LLMs in simulation, we proposed Adaptive Learning Rate Based ORPO (AdaORPO), a reinforcement learning-based algorithm to improve the reliability in simulation across 7 LLMs. Our research provides a foundation for future studies to explore more robust and trustworthy LLM-based simulations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23426
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?
Huang, Yue
Yuan, Zhengqing
Zhou, Yujun
Guo, Kehan
Wang, Xiangqi
Zhuang, Haomin
Sun, Weixiang
Sun, Lichao
Wang, Jindong
Ye, Yanfang
Zhang, Xiangliang
Computation and Language
Large Language Models (LLMs) are increasingly employed for simulations, enabling applications in role-playing agents and Computational Social Science (CSS). However, the reliability of these simulations is under-explored, which raises concerns about the trustworthiness of LLMs in these applications. In this paper, we aim to answer ``How reliable is LLM-based simulation?'' To address this, we introduce TrustSim, an evaluation dataset covering 10 CSS-related topics, to systematically investigate the reliability of the LLM simulation. We conducted experiments on 14 LLMs and found that inconsistencies persist in the LLM-based simulated roles. In addition, the consistency level of LLMs does not strongly correlate with their general performance. To enhance the reliability of LLMs in simulation, we proposed Adaptive Learning Rate Based ORPO (AdaORPO), a reinforcement learning-based algorithm to improve the reliability in simulation across 7 LLMs. Our research provides a foundation for future studies to explore more robust and trustworthy LLM-based simulations.
title Social Science Meets LLMs: How Reliable Are Large Language Models in Social Simulations?
topic Computation and Language
url https://arxiv.org/abs/2410.23426