DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chuang, Yun-Shiuan, Tu, Ruixuan, Dai, Chengtao, Li, You, Vasani, Smit, Yao, Binwei, Tessler, Michael Henry, Yang, Sijia, Shah, Dhavan, Hawkins, Robert, Hu, Junjie, Rogers, Timothy T.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910270571413504
author Chuang, Yun-Shiuan
Tu, Ruixuan
Dai, Chengtao
Li, You
Vasani, Smit
Yao, Binwei
Tessler, Michael Henry
Yang, Sijia
Shah, Dhavan
Hawkins, Robert
Hu, Junjie
Rogers, Timothy T.
author_facet Chuang, Yun-Shiuan
Tu, Ruixuan
Dai, Chengtao
Li, You
Vasani, Smit
Yao, Binwei
Tessler, Michael Henry
Yang, Sijia
Shah, Dhavan
Hawkins, Robert
Hu, Junjie
Rogers, Timothy T.
contents Accurately modeling opinion change through social interactions is crucial for understanding and mitigating polarization, misinformation, and societal conflict. Recent work simulates opinion dynamics with role-playing LLM agents (RPLAs), but multi-agent simulations often display unnatural group behavior, such as premature convergence, and lack empirical benchmarks for assessing alignment with real human group interactions. We introduce DEBATE, a large-scale benchmark for evaluating the authenticity of opinion dynamics in multi-agent RPLA simulations. DEBATE contains multi-round public messages and private Likert-scale beliefs from U.S.-based participants across 107 topics; the cleaned benchmark used in our experiments contains 2,788 participants in 697 groups, enabling evaluation at the utterance and group levels and supporting future individual-level analyses. We instantiate "digital twin" RPLAs with seven LLMs and evaluate across two settings: next-message prediction and full dynamics simulation, using stance-based opinion-dynamics metrics. In zero-shot settings, RPLA groups exhibit strong opinion convergence relative to human groups. On the held-out group split, supervised fine-tuning (SFT) for Llama-3.1-8B-Instruct improves auxiliary stance alignment and reduces group-level convergence error, though discrepancies in opinion change and belief updating remain. DEBATE enables rigorous benchmarking of simulated opinion dynamics and supports future research on aligning multi-agent RPLAs with realistic human interactions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25110
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents
Chuang, Yun-Shiuan
Tu, Ruixuan
Dai, Chengtao
Li, You
Vasani, Smit
Yao, Binwei
Tessler, Michael Henry
Yang, Sijia
Shah, Dhavan
Hawkins, Robert
Hu, Junjie
Rogers, Timothy T.
Computation and Language
Accurately modeling opinion change through social interactions is crucial for understanding and mitigating polarization, misinformation, and societal conflict. Recent work simulates opinion dynamics with role-playing LLM agents (RPLAs), but multi-agent simulations often display unnatural group behavior, such as premature convergence, and lack empirical benchmarks for assessing alignment with real human group interactions. We introduce DEBATE, a large-scale benchmark for evaluating the authenticity of opinion dynamics in multi-agent RPLA simulations. DEBATE contains multi-round public messages and private Likert-scale beliefs from U.S.-based participants across 107 topics; the cleaned benchmark used in our experiments contains 2,788 participants in 697 groups, enabling evaluation at the utterance and group levels and supporting future individual-level analyses. We instantiate "digital twin" RPLAs with seven LLMs and evaluate across two settings: next-message prediction and full dynamics simulation, using stance-based opinion-dynamics metrics. In zero-shot settings, RPLA groups exhibit strong opinion convergence relative to human groups. On the held-out group split, supervised fine-tuning (SFT) for Llama-3.1-8B-Instruct improves auxiliary stance alignment and reduces group-level convergence error, though discrepancies in opinion change and belief updating remain. DEBATE enables rigorous benchmarking of simulated opinion dynamics and supports future research on aligning multi-agent RPLAs with realistic human interactions.
title DEBATE: A Large-Scale Benchmark for Evaluating Opinion Dynamics in Role-Playing LLM Agents
topic Computation and Language
url https://arxiv.org/abs/2510.25110