BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Jiaqi, Lou, Qian, Zheng, Mengxin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916449418739712
author Xue, Jiaqi
Lou, Qian
Zheng, Mengxin
author_facet Xue, Jiaqi
Lou, Qian
Zheng, Mengxin
contents Attacking fairness is crucial because compromised models can introduce biased outcomes, undermining trust and amplifying inequalities in sensitive applications like hiring, healthcare, and law enforcement. This highlights the urgent need to understand how fairness mechanisms can be exploited and to develop defenses that ensure both fairness and robustness. We introduce BadFair, a novel backdoored fairness attack methodology. BadFair stealthily crafts a model that operates with accuracy and fairness under regular conditions but, when activated by certain triggers, discriminates and produces incorrect results for specific groups. This type of attack is particularly stealthy and dangerous, as it circumvents existing fairness detection methods, maintaining an appearance of fairness in normal use. Our findings reveal that BadFair achieves a more than 85% attack success rate in attacks aimed at target groups on average while only incurring a minimal accuracy loss. Moreover, it consistently exhibits a significant discrimination score, distinguishing between pre-defined target and non-target attacked groups across various datasets and models.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17492
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
Xue, Jiaqi
Lou, Qian
Zheng, Mengxin
Cryptography and Security
Computation and Language
Computers and Society
Machine Learning
Attacking fairness is crucial because compromised models can introduce biased outcomes, undermining trust and amplifying inequalities in sensitive applications like hiring, healthcare, and law enforcement. This highlights the urgent need to understand how fairness mechanisms can be exploited and to develop defenses that ensure both fairness and robustness. We introduce BadFair, a novel backdoored fairness attack methodology. BadFair stealthily crafts a model that operates with accuracy and fairness under regular conditions but, when activated by certain triggers, discriminates and produces incorrect results for specific groups. This type of attack is particularly stealthy and dangerous, as it circumvents existing fairness detection methods, maintaining an appearance of fairness in normal use. Our findings reveal that BadFair achieves a more than 85% attack success rate in attacks aimed at target groups on average while only incurring a minimal accuracy loss. Moreover, it consistently exhibits a significant discrimination score, distinguishing between pre-defined target and non-target attacked groups across various datasets and models.
title BadFair: Backdoored Fairness Attacks with Group-conditioned Triggers
topic Cryptography and Security
Computation and Language
Computers and Society
Machine Learning
url https://arxiv.org/abs/2410.17492