Perturbation Effects on Accuracy and Fairness among Similar Individuals

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Xuran, Xue, Hao, Wu, Peng, Ma, Xingjun, Zhang, Zhen, Chen, Huaming, Salim, Flora D.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910274125037568
author Li, Xuran
Xue, Hao
Wu, Peng
Ma, Xingjun
Zhang, Zhen
Chen, Huaming
Salim, Flora D.
author_facet Li, Xuran
Xue, Hao
Wu, Peng
Ma, Xingjun
Zhang, Zhen
Chen, Huaming
Salim, Flora D.
contents Deep neural networks are vulnerable to adversarial perturbations that can simultaneously degrade prediction robustness and individual fairness across diverse application settings. However, existing evaluation protocols typically assess these dimensions in isolation, thereby obscuring critical failure modes. To bridge this gap, we formalize Robust Individual Fairness (RIF): under semantic-preserving (truth-condition-preserving) perturbations, predictions should remain both correct with respect to the ground truth and invariant across semantically equivalent individuals. To surface RIF violations in practice, we introduce RIFair, a black-box adversarial framework that leverages a decoupled perturbation strategy to construct semantically preserved yet unrobust and/or unfair instance pairs. Experiments across multiple model architectures and real-world textual datasets show that robustness-only or fairness-only metrics often miss Robust Biased and Unrobust Fair behaviors. RIFair}reliably exposes these hidden vulnerabilities, supporting RIF as a necessary criterion for trustworthy model assessment. The experimental code is publicly available at https://github.com/Xuran-LI/RIFair.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Perturbation Effects on Accuracy and Fairness among Similar Individuals
Li, Xuran
Xue, Hao
Wu, Peng
Ma, Xingjun
Zhang, Zhen
Chen, Huaming
Salim, Flora D.
Machine Learning
Artificial Intelligence
Computers and Society
Deep neural networks are vulnerable to adversarial perturbations that can simultaneously degrade prediction robustness and individual fairness across diverse application settings. However, existing evaluation protocols typically assess these dimensions in isolation, thereby obscuring critical failure modes. To bridge this gap, we formalize Robust Individual Fairness (RIF): under semantic-preserving (truth-condition-preserving) perturbations, predictions should remain both correct with respect to the ground truth and invariant across semantically equivalent individuals. To surface RIF violations in practice, we introduce RIFair, a black-box adversarial framework that leverages a decoupled perturbation strategy to construct semantically preserved yet unrobust and/or unfair instance pairs. Experiments across multiple model architectures and real-world textual datasets show that robustness-only or fairness-only metrics often miss Robust Biased and Unrobust Fair behaviors. RIFair}reliably exposes these hidden vulnerabilities, supporting RIF as a necessary criterion for trustworthy model assessment. The experimental code is publicly available at https://github.com/Xuran-LI/RIFair.
title Perturbation Effects on Accuracy and Fairness among Similar Individuals
topic Machine Learning
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2404.01356