Guardado en:
Detalles Bibliográficos
Autores principales: Rupprecht, Jens, Ahnert, Georg, Strohmaier, Markus
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2507.07188
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912650671161344
author Rupprecht, Jens
Ahnert, Georg
Strohmaier, Markus
author_facet Rupprecht, Jens
Ahnert, Georg
Strohmaier, Markus
contents Large Language Models (LLMs) are increasingly used as proxies for human subjects in social science surveys, but their reliability and susceptibility to known human-like response biases, such as central tendency, opinion floating and primacy bias are poorly understood. This work investigates the response robustness of LLMs in normative survey contexts, we test nine LLMs on questions from the World Values Survey (WVS), applying a comprehensive set of ten perturbations to both question phrasing and answer option structure, resulting in over 167,000 simulated survey interviews. In doing so, we not only reveal LLMs' vulnerabilities to perturbations but also show that all tested models exhibit a consistent recency bias, disproportionately favoring the last-presented answer option. While larger models are generally more robust, all models remain sensitive to semantic variations like paraphrasing and to combined perturbations. This underscores the critical importance of prompt design and robustness testing when using LLMs to generate synthetic survey data.
format Preprint
id arxiv_https___arxiv_org_abs_2507_07188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses
Rupprecht, Jens
Ahnert, Georg
Strohmaier, Markus
Computation and Language
Artificial Intelligence
Computers and Society
J.4
Large Language Models (LLMs) are increasingly used as proxies for human subjects in social science surveys, but their reliability and susceptibility to known human-like response biases, such as central tendency, opinion floating and primacy bias are poorly understood. This work investigates the response robustness of LLMs in normative survey contexts, we test nine LLMs on questions from the World Values Survey (WVS), applying a comprehensive set of ten perturbations to both question phrasing and answer option structure, resulting in over 167,000 simulated survey interviews. In doing so, we not only reveal LLMs' vulnerabilities to perturbations but also show that all tested models exhibit a consistent recency bias, disproportionately favoring the last-presented answer option. While larger models are generally more robust, all models remain sensitive to semantic variations like paraphrasing and to combined perturbations. This underscores the critical importance of prompt design and robustness testing when using LLMs to generate synthetic survey data.
title Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses
topic Computation and Language
Artificial Intelligence
Computers and Society
J.4
url https://arxiv.org/abs/2507.07188