Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Görge, Rebekka, Gannamaneni, Sujan Sai, Naeven, Tabea, Abdelwahab, Hammam, Allende-Cid, Héctor, Cremers, Armin B., Helmer, Lennard, Mock, Michael, Schmitz, Anna, Xue, Songkai, Yildirir, Elif, Poretschkin, Maximilian, Wrobel, Stefan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915670599401472
author Görge, Rebekka
Gannamaneni, Sujan Sai
Naeven, Tabea
Abdelwahab, Hammam
Allende-Cid, Héctor
Cremers, Armin B.
Helmer, Lennard
Mock, Michael
Schmitz, Anna
Xue, Songkai
Yildirir, Elif
Poretschkin, Maximilian
Wrobel, Stefan
author_facet Görge, Rebekka
Gannamaneni, Sujan Sai
Naeven, Tabea
Abdelwahab, Hammam
Allende-Cid, Héctor
Cremers, Armin B.
Helmer, Lennard
Mock, Michael
Schmitz, Anna
Xue, Songkai
Yildirir, Elif
Poretschkin, Maximilian
Wrobel, Stefan
contents Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations encompassing harmful language and skewed demographic distributions. Regulations such as the European AI Act require identifying and mitigating biases against protected groups in data, with the ultimate goal of preventing unfair model outputs. However, practical guidance and operationalization are lacking. We propose a comprehensive data bias detection and mitigation pipeline comprising four components that address two data bias types, namely representation bias and (explicit) stereotypes for a configurable sensitive attribute. First, we leverage LLM-generated word lists created based on quality criteria to detect relevant group labels. Second, representation bias is quantified using the Demographic Representation Score. Third, we detect and mitigate stereotypes using sociolinguistically informed filtering. Finally, we compensate representation bias through Grammar- and Context-Aware Counterfactual Data Augmentation. We conduct a two-fold evaluation using the examples of gender, religion and age. First, the effectiveness of each individual component on data debiasing is evaluated through human validation and baseline comparison. The findings demonstrate that we successfully reduce representation bias and (explicit) stereotypes in a text dataset. Second, the effect of data debiasing on model bias reduction is evaluated by bias benchmarking of several models (0.6B-8B parameters), fine-tuned on the debiased text dataset. This evaluation reveals that LLMs fine-tuned on debiased data do not consistently show improved performance on bias benchmarks, exposing critical gaps in current evaluation methodologies and highlighting the need for targeted data manipulation to address manifested model bias.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10734
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
Görge, Rebekka
Gannamaneni, Sujan Sai
Naeven, Tabea
Abdelwahab, Hammam
Allende-Cid, Héctor
Cremers, Armin B.
Helmer, Lennard
Mock, Michael
Schmitz, Anna
Xue, Songkai
Yildirir, Elif
Poretschkin, Maximilian
Wrobel, Stefan
Computation and Language
Artificial Intelligence
68T50
I.2; I.2.7
Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations encompassing harmful language and skewed demographic distributions. Regulations such as the European AI Act require identifying and mitigating biases against protected groups in data, with the ultimate goal of preventing unfair model outputs. However, practical guidance and operationalization are lacking. We propose a comprehensive data bias detection and mitigation pipeline comprising four components that address two data bias types, namely representation bias and (explicit) stereotypes for a configurable sensitive attribute. First, we leverage LLM-generated word lists created based on quality criteria to detect relevant group labels. Second, representation bias is quantified using the Demographic Representation Score. Third, we detect and mitigate stereotypes using sociolinguistically informed filtering. Finally, we compensate representation bias through Grammar- and Context-Aware Counterfactual Data Augmentation. We conduct a two-fold evaluation using the examples of gender, religion and age. First, the effectiveness of each individual component on data debiasing is evaluated through human validation and baseline comparison. The findings demonstrate that we successfully reduce representation bias and (explicit) stereotypes in a text dataset. Second, the effect of data debiasing on model bias reduction is evaluated by bias benchmarking of several models (0.6B-8B parameters), fine-tuned on the debiased text dataset. This evaluation reveals that LLMs fine-tuned on debiased data do not consistently show improved performance on bias benchmarks, exposing critical gaps in current evaluation methodologies and highlighting the need for targeted data manipulation to address manifested model bias.
title Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
topic Computation and Language
Artificial Intelligence
68T50
I.2; I.2.7
url https://arxiv.org/abs/2512.10734