Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mooney, James, Woldense, Josef, Jia, Zheng Robert, Hayati, Shirley Anugrah, Nguyen, My Ha, Raheja, Vipul, Kang, Dongyeop
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913103331983360
author Mooney, James
Woldense, Josef
Jia, Zheng Robert
Hayati, Shirley Anugrah
Nguyen, My Ha
Raheja, Vipul
Kang, Dongyeop
author_facet Mooney, James
Woldense, Josef
Jia, Zheng Robert
Hayati, Shirley Anugrah
Nguyen, My Ha
Raheja, Vipul
Kang, Dongyeop
contents The impressive capabilities of Large Language Models (LLMs) raise the possibility that synthetic agents can serve as substitutes for real participants in human-subject research. To evaluate this claim, prior research has largely focused on whether LLM-generated survey responses align with those produced by human respondents whom the LLMs are prompted to represent. In contrast, we address a more fundamental question: Do agents maintain empirical consistency; aligning to human behavioral models when examined under different experimental settings? To this end, we develop a study designed to (a) ask a set of questions which reveals an agent's latent profile and (b) examine agent behavioral consistency in a conversational setting with other agents. This design enables us to explore a set of behavioral hypotheses to assess whether an agent's conversational behavior is consistent with what we would expect from its revealed state. Our findings show significant inconsistencies in LLMs across model families and at differing model sizes. Most importantly, we find that, although agents may generate responses matching those of their human counterparts, they fail to be empirically consistent, representing a critical gap in their capabilities to accurately substitute for real participants in human-subject research.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03736
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
Mooney, James
Woldense, Josef
Jia, Zheng Robert
Hayati, Shirley Anugrah
Nguyen, My Ha
Raheja, Vipul
Kang, Dongyeop
Artificial Intelligence
Computation and Language
Machine Learning
The impressive capabilities of Large Language Models (LLMs) raise the possibility that synthetic agents can serve as substitutes for real participants in human-subject research. To evaluate this claim, prior research has largely focused on whether LLM-generated survey responses align with those produced by human respondents whom the LLMs are prompted to represent. In contrast, we address a more fundamental question: Do agents maintain empirical consistency; aligning to human behavioral models when examined under different experimental settings? To this end, we develop a study designed to (a) ask a set of questions which reveals an agent's latent profile and (b) examine agent behavioral consistency in a conversational setting with other agents. This design enables us to explore a set of behavioral hypotheses to assess whether an agent's conversational behavior is consistent with what we would expect from its revealed state. Our findings show significant inconsistencies in LLMs across model families and at differing model sizes. Most importantly, we find that, although agents may generate responses matching those of their human counterparts, they fail to be empirically consistent, representing a critical gap in their capabilities to accurately substitute for real participants in human-subject research.
title Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.03736