Utilizing Adversarial Examples for Bias Mitigation and Accuracy Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shukla, Pushkar, Srikanth, Dhruv, Cohen, Lee, Turk, Matthew
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916304639754240
author Shukla, Pushkar
Srikanth, Dhruv
Cohen, Lee
Turk, Matthew
author_facet Shukla, Pushkar
Srikanth, Dhruv
Cohen, Lee
Turk, Matthew
contents We propose a novel approach to mitigate biases in computer vision models by utilizing counterfactual generation and fine-tuning. While counterfactuals have been used to analyze and address biases in DNN models, the counterfactuals themselves are often generated from biased generative models, which can introduce additional biases or spurious correlations. To address this issue, we propose using adversarial images, that is images that deceive a deep neural network but not humans, as counterfactuals for fair model training. Our approach leverages a curriculum learning framework combined with a fine-grained adversarial loss to fine-tune the model using adversarial examples. By incorporating adversarial images into the training data, we aim to prevent biases from propagating through the pipeline. We validate our approach through both qualitative and quantitative assessments, demonstrating improved bias mitigation and accuracy compared to existing methods. Qualitatively, our results indicate that post-training, the decisions made by the model are less dependent on the sensitive attribute and our model better disentangles the relationship between sensitive attributes and classification variables.
format Preprint
id arxiv_https___arxiv_org_abs_2404_11819
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Utilizing Adversarial Examples for Bias Mitigation and Accuracy Enhancement
Shukla, Pushkar
Srikanth, Dhruv
Cohen, Lee
Turk, Matthew
Computer Vision and Pattern Recognition
We propose a novel approach to mitigate biases in computer vision models by utilizing counterfactual generation and fine-tuning. While counterfactuals have been used to analyze and address biases in DNN models, the counterfactuals themselves are often generated from biased generative models, which can introduce additional biases or spurious correlations. To address this issue, we propose using adversarial images, that is images that deceive a deep neural network but not humans, as counterfactuals for fair model training. Our approach leverages a curriculum learning framework combined with a fine-grained adversarial loss to fine-tune the model using adversarial examples. By incorporating adversarial images into the training data, we aim to prevent biases from propagating through the pipeline. We validate our approach through both qualitative and quantitative assessments, demonstrating improved bias mitigation and accuracy compared to existing methods. Qualitatively, our results indicate that post-training, the decisions made by the model are less dependent on the sensitive attribute and our model better disentangles the relationship between sensitive attributes and classification variables.
title Utilizing Adversarial Examples for Bias Mitigation and Accuracy Enhancement
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2404.11819