Benchmarking Gender and Political Bias in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jinrui, Han, Xudong, Baldwin, Timothy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914038997319680
author Yang, Jinrui
Han, Xudong
Baldwin, Timothy
author_facet Yang, Jinrui
Han, Xudong
Baldwin, Timothy
contents We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-call vote outcomes and includes rich demographic metadata for each Member of the European Parliament (MEP), such as gender, age, country, and political group. Using EuroParlVote, we evaluate state-of-the-art LLMs on two tasks -- gender classification and vote prediction -- revealing consistent patterns of bias. We find that LLMs frequently misclassify female MEPs as male and demonstrate reduced accuracy when simulating votes for female speakers. Politically, LLMs tend to favor centrist groups while underperforming on both far-left and far-right ones. Proprietary models like GPT-4o outperform open-weight alternatives in terms of both robustness and fairness. We release the EuroParlVote dataset, code, and demo to support future research on fairness and accountability in NLP within political contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06164
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Gender and Political Bias in Large Language Models
Yang, Jinrui
Han, Xudong
Baldwin, Timothy
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
We introduce EuroParlVote, a novel benchmark for evaluating large language models (LLMs) in politically sensitive contexts. It links European Parliament debate speeches to roll-call vote outcomes and includes rich demographic metadata for each Member of the European Parliament (MEP), such as gender, age, country, and political group. Using EuroParlVote, we evaluate state-of-the-art LLMs on two tasks -- gender classification and vote prediction -- revealing consistent patterns of bias. We find that LLMs frequently misclassify female MEPs as male and demonstrate reduced accuracy when simulating votes for female speakers. Politically, LLMs tend to favor centrist groups while underperforming on both far-left and far-right ones. Proprietary models like GPT-4o outperform open-weight alternatives in terms of both robustness and fairness. We release the EuroParlVote dataset, code, and demo to support future research on fairness and accountability in NLP within political contexts.
title Benchmarking Gender and Political Bias in Large Language Models
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2509.06164