Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tomar, Aditya, Murthy, Rudra, Bhattacharyya, Pushpak
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909673328738304
author Tomar, Aditya
Murthy, Rudra
Bhattacharyya, Pushpak
author_facet Tomar, Aditya
Murthy, Rudra
Bhattacharyya, Pushpak
contents Bias and stereotypes in language models can cause harm, especially in sensitive areas like content moderation and decision-making. This paper addresses bias and stereotype detection by exploring how jointly learning these tasks enhances model performance. We introduce StereoBias, a unique dataset labeled for bias and stereotype detection across five categories: religion, gender, socio-economic status, race, profession, and others, enabling a deeper study of their relationship. Our experiments compare encoder-only models and fine-tuned decoder-only models using QLoRA. While encoder-only models perform well, decoder-only models also show competitive results. Crucially, joint training on bias and stereotype detection significantly improves bias detection compared to training them separately. Additional experiments with sentiment analysis confirm that the improvements stem from the connection between bias and stereotypes, not multi-task learning alone. These findings highlight the value of leveraging stereotype information to build fairer and more effective AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01715
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
Tomar, Aditya
Murthy, Rudra
Bhattacharyya, Pushpak
Computation and Language
Bias and stereotypes in language models can cause harm, especially in sensitive areas like content moderation and decision-making. This paper addresses bias and stereotype detection by exploring how jointly learning these tasks enhances model performance. We introduce StereoBias, a unique dataset labeled for bias and stereotype detection across five categories: religion, gender, socio-economic status, race, profession, and others, enabling a deeper study of their relationship. Our experiments compare encoder-only models and fine-tuned decoder-only models using QLoRA. While encoder-only models perform well, decoder-only models also show competitive results. Crucially, joint training on bias and stereotype detection significantly improves bias detection compared to training them separately. Additional experiments with sentiment analysis confirm that the improvements stem from the connection between bias and stereotypes, not multi-task learning alone. These findings highlight the value of leveraging stereotype information to build fairer and more effective AI systems.
title Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
topic Computation and Language
url https://arxiv.org/abs/2507.01715