Digital Twins as Funhouse Mirrors: Five Key Distortions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Peng, Tianyi, Gui, George, Brucks, Melanie, Merlau, Daniel J., Fan, Grace Jiarui, Sliman, Malek Ben, Johnson, Eric J., Althenayyan, Abdullah, Bellezza, Silvia, Donati, Dante, Fong, Hortense, Friedman, Elizabeth, Guevara, Ariana, Hussein, Mohamed, Jerath, Kinshuk, Kogut, Bruce, Kumar, Akshit, Lane, Kristen, Li, Hannah, Morwitz, Vicki, Netzer, Oded, Perkowski, Patryk, Toubia, Olivier
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911605117157376
author Peng, Tianyi
Gui, George
Brucks, Melanie
Merlau, Daniel J.
Fan, Grace Jiarui
Sliman, Malek Ben
Johnson, Eric J.
Althenayyan, Abdullah
Bellezza, Silvia
Donati, Dante
Fong, Hortense
Friedman, Elizabeth
Guevara, Ariana
Hussein, Mohamed
Jerath, Kinshuk
Kogut, Bruce
Kumar, Akshit
Lane, Kristen
Li, Hannah
Morwitz, Vicki
Netzer, Oded
Perkowski, Patryk
Toubia, Olivier
author_facet Peng, Tianyi
Gui, George
Brucks, Melanie
Merlau, Daniel J.
Fan, Grace Jiarui
Sliman, Malek Ben
Johnson, Eric J.
Althenayyan, Abdullah
Bellezza, Silvia
Donati, Dante
Fong, Hortense
Friedman, Elizabeth
Guevara, Ariana
Hussein, Mohamed
Jerath, Kinshuk
Kogut, Bruce
Kumar, Akshit
Lane, Kristen
Li, Hannah
Morwitz, Vicki
Netzer, Oded
Perkowski, Patryk
Toubia, Olivier
contents Scientists and practitioners are increasingly moving to deploy digital twins--LLM-based models of real individuals--across social science and policy research. We conduct 19 pre-registered studies spanning 164 diverse outcomes (e.g., attitudes toward hiring algorithms, intentions to share misinformation), comparing human responses to those of their corresponding digital twins, which are trained on each individual's prior responses to over 500 questions. We establish an empirical benchmark for digital twin performance: their predictions are only modestly more accurate than those of a homogeneous base LLM and exhibit weak correlation with human responses (average $r = 0.20$). To inform future development, we identify five systematic distortions in digital twin behavior: (i) insufficient individuation, (ii) stereotyping, (iii) representation bias, (iv) ideological bias, and (v) hyper-rationality. Finally, we release our full dataset and code as a standardized testbed for evaluating and improving digital twin methodologies. Together, our findings caution against premature deployment while laying the groundwork for a transparent, replicable, and iterative science of responsible digital twin development.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19088
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Digital Twins as Funhouse Mirrors: Five Key Distortions
Peng, Tianyi
Gui, George
Brucks, Melanie
Merlau, Daniel J.
Fan, Grace Jiarui
Sliman, Malek Ben
Johnson, Eric J.
Althenayyan, Abdullah
Bellezza, Silvia
Donati, Dante
Fong, Hortense
Friedman, Elizabeth
Guevara, Ariana
Hussein, Mohamed
Jerath, Kinshuk
Kogut, Bruce
Kumar, Akshit
Lane, Kristen
Li, Hannah
Morwitz, Vicki
Netzer, Oded
Perkowski, Patryk
Toubia, Olivier
Computers and Society
Artificial Intelligence
Human-Computer Interaction
Applications
Scientists and practitioners are increasingly moving to deploy digital twins--LLM-based models of real individuals--across social science and policy research. We conduct 19 pre-registered studies spanning 164 diverse outcomes (e.g., attitudes toward hiring algorithms, intentions to share misinformation), comparing human responses to those of their corresponding digital twins, which are trained on each individual's prior responses to over 500 questions. We establish an empirical benchmark for digital twin performance: their predictions are only modestly more accurate than those of a homogeneous base LLM and exhibit weak correlation with human responses (average $r = 0.20$). To inform future development, we identify five systematic distortions in digital twin behavior: (i) insufficient individuation, (ii) stereotyping, (iii) representation bias, (iv) ideological bias, and (v) hyper-rationality. Finally, we release our full dataset and code as a standardized testbed for evaluating and improving digital twin methodologies. Together, our findings caution against premature deployment while laying the groundwork for a transparent, replicable, and iterative science of responsible digital twin development.
title Digital Twins as Funhouse Mirrors: Five Key Distortions
topic Computers and Society
Artificial Intelligence
Human-Computer Interaction
Applications
url https://arxiv.org/abs/2509.19088