Does Differential Privacy Impact Bias in Pretrained NLP Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Islam, Md. Khairul, Wang, Andrew, Wang, Tianhao, Ji, Yangfeng, Fox, Judy, Zhao, Jieyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910664785657856
author Islam, Md. Khairul
Wang, Andrew
Wang, Tianhao
Ji, Yangfeng
Fox, Judy
Zhao, Jieyu
author_facet Islam, Md. Khairul
Wang, Andrew
Wang, Tianhao
Ji, Yangfeng
Fox, Judy
Zhao, Jieyu
contents Differential privacy (DP) is applied when fine-tuning pre-trained large language models (LLMs) to limit leakage of training examples. While most DP research has focused on improving a model's privacy-utility tradeoff, some find that DP can be unfair to or biased against underrepresented groups. In this work, we show the impact of DP on bias in LLMs through empirical analysis. Differentially private training can increase the model bias against protected groups w.r.t AUC-based bias metrics. DP makes it more difficult for the model to differentiate between the positive and negative examples from the protected groups and other groups in the rest of the population. Our results also show that the impact of DP on bias is not only affected by the privacy protection level but also the underlying distribution of the dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2410_18749
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Does Differential Privacy Impact Bias in Pretrained NLP Models?
Islam, Md. Khairul
Wang, Andrew
Wang, Tianhao
Ji, Yangfeng
Fox, Judy
Zhao, Jieyu
Computation and Language
Artificial Intelligence
Machine Learning
Differential privacy (DP) is applied when fine-tuning pre-trained large language models (LLMs) to limit leakage of training examples. While most DP research has focused on improving a model's privacy-utility tradeoff, some find that DP can be unfair to or biased against underrepresented groups. In this work, we show the impact of DP on bias in LLMs through empirical analysis. Differentially private training can increase the model bias against protected groups w.r.t AUC-based bias metrics. DP makes it more difficult for the model to differentiate between the positive and negative examples from the protected groups and other groups in the rest of the population. Our results also show that the impact of DP on bias is not only affected by the privacy protection level but also the underlying distribution of the dataset.
title Does Differential Privacy Impact Bias in Pretrained NLP Models?
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.18749