High Risk of Political Bias in Black Box Emotion Inference Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Plisiecki, Hubert, Lenartowicz, Paweł, Flakus, Maria, Pokropek, Artur
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909398002040832
author Plisiecki, Hubert
Lenartowicz, Paweł
Flakus, Maria
Pokropek, Artur
author_facet Plisiecki, Hubert
Lenartowicz, Paweł
Flakus, Maria
Pokropek, Artur
contents This paper investigates the presence of political bias in emotion inference models used for sentiment analysis (SA) in social science research. Machine learning models often reflect biases in their training data, impacting the validity of their outcomes. While previous research has highlighted gender and race biases, our study focuses on political bias - an underexplored yet pervasive issue that can skew the interpretation of text data across a wide array of studies. We conducted a bias audit on a Polish sentiment analysis model developed in our lab. By analyzing valence predictions for names and sentences involving Polish politicians, we uncovered systematic differences influenced by political affiliations. Our findings indicate that annotations by human raters propagate political biases into the model's predictions. To mitigate this, we pruned the training dataset of texts mentioning these politicians and observed a reduction in bias, though not its complete elimination. Given the significant implications of political bias in SA, our study emphasizes caution in employing these models for social science research. We recommend a critical examination of SA results and propose using lexicon-based systems as a more ideologically neutral alternative. This paper underscores the necessity for ongoing scrutiny and methodological adjustments to ensure the reliability and impartiality of the use of machine learning in academic and applied contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2407_13891
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle High Risk of Political Bias in Black Box Emotion Inference Models
Plisiecki, Hubert
Lenartowicz, Paweł
Flakus, Maria
Pokropek, Artur
Computation and Language
Artificial Intelligence
This paper investigates the presence of political bias in emotion inference models used for sentiment analysis (SA) in social science research. Machine learning models often reflect biases in their training data, impacting the validity of their outcomes. While previous research has highlighted gender and race biases, our study focuses on political bias - an underexplored yet pervasive issue that can skew the interpretation of text data across a wide array of studies. We conducted a bias audit on a Polish sentiment analysis model developed in our lab. By analyzing valence predictions for names and sentences involving Polish politicians, we uncovered systematic differences influenced by political affiliations. Our findings indicate that annotations by human raters propagate political biases into the model's predictions. To mitigate this, we pruned the training dataset of texts mentioning these politicians and observed a reduction in bias, though not its complete elimination. Given the significant implications of political bias in SA, our study emphasizes caution in employing these models for social science research. We recommend a critical examination of SA results and propose using lexicon-based systems as a more ideologically neutral alternative. This paper underscores the necessity for ongoing scrutiny and methodological adjustments to ensure the reliability and impartiality of the use of machine learning in academic and applied contexts.
title High Risk of Political Bias in Black Box Emotion Inference Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.13891