When Can Digital Personas Reliably Approximate Human Survey Findings?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jia, Mumin, Chen, Yilin, Sharma, Divya, Diaz-Rodriguez, Jairo
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913112737710080
author Jia, Mumin
Chen, Yilin
Sharma, Divya
Diaz-Rodriguez, Jairo
author_facet Jia, Mumin
Chen, Yilin
Sharma, Divya
Diaz-Rodriguez, Jairo
contents Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.
format Preprint
id arxiv_https___arxiv_org_abs_2605_10659
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Can Digital Personas Reliably Approximate Human Survey Findings?
Jia, Mumin
Chen, Yilin
Sharma, Divya
Diaz-Rodriguez, Jairo
Computation and Language
Artificial Intelligence
Social and Information Networks
Machine Learning
Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.
title When Can Digital Personas Reliably Approximate Human Survey Findings?
topic Computation and Language
Artificial Intelligence
Social and Information Networks
Machine Learning
url https://arxiv.org/abs/2605.10659