Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Dhaini, Mahdi, Erdogan, Ege, Feldhus, Nils, Kasneci, Gjergji
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909598902910976
author Dhaini, Mahdi
Erdogan, Ege
Feldhus, Nils
Kasneci, Gjergji
author_facet Dhaini, Mahdi
Erdogan, Ege
Feldhus, Nils
Kasneci, Gjergji
contents While research on applications and evaluations of explanation methods continues to expand, fairness of the explanation methods concerning disparities in their performance across subgroups remains an often overlooked aspect. In this paper, we address this gap by showing that, across three tasks and five language models, widely used post-hoc feature attribution methods exhibit significant gender disparity with respect to their faithfulness, robustness, and complexity. These disparities persist even when the models are pre-trained or fine-tuned on particularly unbiased datasets, indicating that the disparities we observe are not merely consequences of biased training data. Our results highlight the importance of addressing disparities in explanations when developing and applying explainability methods, as these can lead to biased outcomes against certain subgroups, with particularly critical implications in high-stakes contexts. Furthermore, our findings underscore the importance of incorporating the fairness of explanations, alongside overall model fairness and explainability, as a requirement in regulatory frameworks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_01198
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
Dhaini, Mahdi
Erdogan, Ege
Feldhus, Nils
Kasneci, Gjergji
Computation and Language
Artificial Intelligence
Machine Learning
While research on applications and evaluations of explanation methods continues to expand, fairness of the explanation methods concerning disparities in their performance across subgroups remains an often overlooked aspect. In this paper, we address this gap by showing that, across three tasks and five language models, widely used post-hoc feature attribution methods exhibit significant gender disparity with respect to their faithfulness, robustness, and complexity. These disparities persist even when the models are pre-trained or fine-tuned on particularly unbiased datasets, indicating that the disparities we observe are not merely consequences of biased training data. Our results highlight the importance of addressing disparities in explanations when developing and applying explainability methods, as these can lead to biased outcomes against certain subgroups, with particularly critical implications in high-stakes contexts. Furthermore, our findings underscore the importance of incorporating the fairness of explanations, alongside overall model fairness and explainability, as a requirement in regulatory frameworks.
title Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.01198