Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ardebili, Ali Aghazadeh, Stella, Massimo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913076339539968
author Ardebili, Ali Aghazadeh
Stella, Massimo
author_facet Ardebili, Ali Aghazadeh
Stella, Massimo
contents Large Language Models (LLMs) can strongly shape social discourse, yet datasets investigating how LLM outputs vary across controlled social and contextual prompting remain sparse. Cognitive Digital Shadows (CDS) is a 190,000-record synthetic corpus supporting analyses of LLM-generated discourse. Each CDS record is generated by one of 19 LLMs, prompted to shadow either a human persona or an AI-assistant role. CDS contains LLM responses on 4 controversial societal topics: vaccines/healthcare, social media disinformation, the gender gap in science, and STEM stereotypes. Persona-conditioned records encode 17 sociodemographic and psychological attributes, providing data linking LLMs' prompts, language, stances and reasoning. Texts are validated for topic anchoring and can support emotional analyses via interpretable NLP (e.g. textual forma mentis networks). CDS is enriched by a pooling platform with user-friendly dashboards, enabling easy, interactive group-level comparisons of emotional and semantic framing across personas, topics and models. The CDS prompting framework supports future audits of LLMs' bias, social sensitivity and alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27624
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
Ardebili, Ali Aghazadeh
Stella, Massimo
Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Machine Learning
Large Language Models (LLMs) can strongly shape social discourse, yet datasets investigating how LLM outputs vary across controlled social and contextual prompting remain sparse. Cognitive Digital Shadows (CDS) is a 190,000-record synthetic corpus supporting analyses of LLM-generated discourse. Each CDS record is generated by one of 19 LLMs, prompted to shadow either a human persona or an AI-assistant role. CDS contains LLM responses on 4 controversial societal topics: vaccines/healthcare, social media disinformation, the gender gap in science, and STEM stereotypes. Persona-conditioned records encode 17 sociodemographic and psychological attributes, providing data linking LLMs' prompts, language, stances and reasoning. Texts are validated for topic anchoring and can support emotional analyses via interpretable NLP (e.g. textual forma mentis networks). CDS is enriched by a pooling platform with user-friendly dashboards, enabling easy, interactive group-level comparisons of emotional and semantic framing across personas, topics and models. The CDS prompting framework supports future audits of LLMs' bias, social sensitivity and alignment.
title Mapping how LLMs debate societal issues when shadowing human personality traits, sociodemographics and social media behavior
topic Computation and Language
Artificial Intelligence
Computers and Society
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2604.27624