Debiasify: Self-Distillation for Unsupervised Bias Mitigation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bayasi, Nourhan, Fayyad, Jamil, Hamarneh, Ghassan, Garbi, Rafeef, Najjaran, Homayoun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917825841463296
author Bayasi, Nourhan
Fayyad, Jamil
Hamarneh, Ghassan
Garbi, Rafeef
Najjaran, Homayoun
author_facet Bayasi, Nourhan
Fayyad, Jamil
Hamarneh, Ghassan
Garbi, Rafeef
Najjaran, Homayoun
contents Simplicity bias poses a significant challenge in neural networks, often leading models to favor simpler solutions and inadvertently learn decision rules influenced by spurious correlations. This results in biased models with diminished generalizability. While many current approaches depend on human supervision, obtaining annotations for various bias attributes is often impractical. To address this, we introduce Debiasify, a novel self-distillation approach that requires no prior knowledge about the nature of biases. Our method leverages a new distillation loss to transfer knowledge within the network, from deeper layers containing complex, highly-predictive features to shallower layers with simpler, attribute-conditioned features in an unsupervised manner. This enables Debiasify to learn robust, debiased representations that generalize effectively across diverse biases and datasets, improving both worst-group performance and overall accuracy. Extensive experiments on computer vision and medical imaging benchmarks demonstrate the effectiveness of our approach, significantly outperforming previous unsupervised debiasing methods (e.g., a 10.13% improvement in worst-group accuracy for Wavy Hair classification in CelebA) and achieving comparable or superior performance to supervised approaches. Our code is publicly available at the following link: Debiasify.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00711
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Debiasify: Self-Distillation for Unsupervised Bias Mitigation
Bayasi, Nourhan
Fayyad, Jamil
Hamarneh, Ghassan
Garbi, Rafeef
Najjaran, Homayoun
Computer Vision and Pattern Recognition
Machine Learning
Simplicity bias poses a significant challenge in neural networks, often leading models to favor simpler solutions and inadvertently learn decision rules influenced by spurious correlations. This results in biased models with diminished generalizability. While many current approaches depend on human supervision, obtaining annotations for various bias attributes is often impractical. To address this, we introduce Debiasify, a novel self-distillation approach that requires no prior knowledge about the nature of biases. Our method leverages a new distillation loss to transfer knowledge within the network, from deeper layers containing complex, highly-predictive features to shallower layers with simpler, attribute-conditioned features in an unsupervised manner. This enables Debiasify to learn robust, debiased representations that generalize effectively across diverse biases and datasets, improving both worst-group performance and overall accuracy. Extensive experiments on computer vision and medical imaging benchmarks demonstrate the effectiveness of our approach, significantly outperforming previous unsupervised debiasing methods (e.g., a 10.13% improvement in worst-group accuracy for Wavy Hair classification in CelebA) and achieving comparable or superior performance to supervised approaches. Our code is publicly available at the following link: Debiasify.
title Debiasify: Self-Distillation for Unsupervised Bias Mitigation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2411.00711