United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: von der Heyde, Leah, Haensch, Anna-Carolina, Wenz, Alexander, Ma, Bolei
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908325170380800
author von der Heyde, Leah
Haensch, Anna-Carolina
Wenz, Alexander
Ma, Bolei
author_facet von der Heyde, Leah
Haensch, Anna-Carolina
Wenz, Alexander
Ma, Bolei
contents "Synthetic samples" based on large language models (LLMs) have been argued to serve as efficient alternatives to surveys of humans, assuming that their training data includes information on human attitudes and behavior. However, LLM-synthetic samples might exhibit bias, for example due to training data and fine-tuning processes being unrepresentative of diverse contexts. Such biases risk reinforcing existing biases in research, policymaking, and society. Therefore, researchers need to investigate if and under which conditions LLM-generated synthetic samples can be used for public opinion prediction. In this study, we examine to what extent LLM-based predictions of individual public opinion exhibit context-dependent biases by predicting the results of the 2024 European Parliament elections. Prompting three LLMs with individual-level background information of 26,000 eligible European voters, we ask the LLMs to predict each person's voting behavior. By comparing them to the actual results, we show that LLM-based predictions of future voting behavior largely fail, their accuracy is unequally distributed across national and linguistic contexts, and they require detailed attitudinal information in the prompt. The findings emphasize the limited applicability of LLM-synthetic samples to public opinion prediction. In investigating their contextual biases, this study contributes to the understanding and mitigation of inequalities in the development of LLMs and their applications in computational social science.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09045
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
von der Heyde, Leah
Haensch, Anna-Carolina
Wenz, Alexander
Ma, Bolei
Computers and Society
Artificial Intelligence
Computation and Language
Applications
"Synthetic samples" based on large language models (LLMs) have been argued to serve as efficient alternatives to surveys of humans, assuming that their training data includes information on human attitudes and behavior. However, LLM-synthetic samples might exhibit bias, for example due to training data and fine-tuning processes being unrepresentative of diverse contexts. Such biases risk reinforcing existing biases in research, policymaking, and society. Therefore, researchers need to investigate if and under which conditions LLM-generated synthetic samples can be used for public opinion prediction. In this study, we examine to what extent LLM-based predictions of individual public opinion exhibit context-dependent biases by predicting the results of the 2024 European Parliament elections. Prompting three LLMs with individual-level background information of 26,000 eligible European voters, we ask the LLMs to predict each person's voting behavior. By comparing them to the actual results, we show that LLM-based predictions of future voting behavior largely fail, their accuracy is unequally distributed across national and linguistic contexts, and they require detailed attitudinal information in the prompt. The findings emphasize the limited applicability of LLM-synthetic samples to public opinion prediction. In investigating their contextual biases, this study contributes to the understanding and mitigation of inequalities in the development of LLMs and their applications in computational social science.
title United in Diversity? Contextual Biases in LLM-Based Predictions of the 2024 European Parliament Elections
topic Computers and Society
Artificial Intelligence
Computation and Language
Applications
url https://arxiv.org/abs/2409.09045