Structured Reasoning for Fairness: A Multi-Agent Approach to Bias Detection in Textual Data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Tianyi, Fan, Elsa
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912253388783616
author Huang, Tianyi
Fan, Elsa
author_facet Huang, Tianyi
Fan, Elsa
contents From disinformation spread by AI chatbots to AI recommendations that inadvertently reinforce stereotypes, textual bias poses a significant challenge to the trustworthiness of large language models (LLMs). In this paper, we propose a multi-agent framework that systematically identifies biases by disentangling each statement as fact or opinion, assigning a bias intensity score, and providing concise, factual justifications. Evaluated on 1,500 samples from the WikiNPOV dataset, the framework achieves 84.9% accuracy$\unicode{x2014}$an improvement of 13.0% over the zero-shot baseline$\unicode{x2014}$demonstrating the efficacy of explicitly modeling fact versus opinion prior to quantifying bias intensity. By combining enhanced detection accuracy with interpretable explanations, this approach sets a foundation for promoting fairness and accountability in modern language models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00355
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structured Reasoning for Fairness: A Multi-Agent Approach to Bias Detection in Textual Data
Huang, Tianyi
Fan, Elsa
Computation and Language
Artificial Intelligence
From disinformation spread by AI chatbots to AI recommendations that inadvertently reinforce stereotypes, textual bias poses a significant challenge to the trustworthiness of large language models (LLMs). In this paper, we propose a multi-agent framework that systematically identifies biases by disentangling each statement as fact or opinion, assigning a bias intensity score, and providing concise, factual justifications. Evaluated on 1,500 samples from the WikiNPOV dataset, the framework achieves 84.9% accuracy$\unicode{x2014}$an improvement of 13.0% over the zero-shot baseline$\unicode{x2014}$demonstrating the efficacy of explicitly modeling fact versus opinion prior to quantifying bias intensity. By combining enhanced detection accuracy with interpretable explanations, this approach sets a foundation for promoting fairness and accountability in modern language models.
title Structured Reasoning for Fairness: A Multi-Agent Approach to Bias Detection in Textual Data
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2503.00355