NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Greco, Salvatore, Zhou, Ke, Capra, Licia, Cerquitelli, Tania, Quercia, Daniele
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916484574347264
author Greco, Salvatore
Zhou, Ke
Capra, Licia
Cerquitelli, Tania
Quercia, Daniele
author_facet Greco, Salvatore
Zhou, Ke
Capra, Licia
Cerquitelli, Tania
Quercia, Daniele
contents AI regulations are expected to prohibit machine learning models from using sensitive attributes during training. However, the latest Natural Language Processing (NLP) classifiers, which rely on deep learning, operate as black-box systems, complicating the detection and remediation of such misuse. Traditional bias mitigation methods in NLP aim for comparable performance across different groups based on attributes like gender or race but fail to address the underlying issue of reliance on protected attributes. To partly fix that, we introduce NLPGuard, a framework for mitigating the reliance on protected attributes in NLP classifiers. NLPGuard takes an unlabeled dataset, an existing NLP classifier, and its training data as input, producing a modified training dataset that significantly reduces dependence on protected attributes without compromising accuracy. NLPGuard is applied to three classification tasks: identifying toxic language, sentiment analysis, and occupation classification. Our evaluation shows that current NLP classifiers heavily depend on protected attributes, with up to $23\%$ of the most predictive words associated with these attributes. However, NLPGuard effectively reduces this reliance by up to $79\%$, while slightly improving accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2407_01697
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers
Greco, Salvatore
Zhou, Ke
Capra, Licia
Cerquitelli, Tania
Quercia, Daniele
Computation and Language
Artificial Intelligence
Human-Computer Interaction
AI regulations are expected to prohibit machine learning models from using sensitive attributes during training. However, the latest Natural Language Processing (NLP) classifiers, which rely on deep learning, operate as black-box systems, complicating the detection and remediation of such misuse. Traditional bias mitigation methods in NLP aim for comparable performance across different groups based on attributes like gender or race but fail to address the underlying issue of reliance on protected attributes. To partly fix that, we introduce NLPGuard, a framework for mitigating the reliance on protected attributes in NLP classifiers. NLPGuard takes an unlabeled dataset, an existing NLP classifier, and its training data as input, producing a modified training dataset that significantly reduces dependence on protected attributes without compromising accuracy. NLPGuard is applied to three classification tasks: identifying toxic language, sentiment analysis, and occupation classification. Our evaluation shows that current NLP classifiers heavily depend on protected attributes, with up to $23\%$ of the most predictive words associated with these attributes. However, NLPGuard effectively reduces this reliance by up to $79\%$, while slightly improving accuracy.
title NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP Classifiers
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2407.01697