Inference-Time Rule Eraser: Fair Recognition via Distilling and Removing Biased Rules

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yi, Lu, Dongyuan, Sang, Jitao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913481530277888
author Zhang, Yi
Lu, Dongyuan
Sang, Jitao
author_facet Zhang, Yi
Lu, Dongyuan
Sang, Jitao
contents Machine learning models often make predictions based on biased features such as gender, race, and other social attributes, posing significant fairness risks, especially in societal applications, such as hiring, banking, and criminal justice. Traditional approaches to addressing this issue involve retraining or fine-tuning neural networks with fairness-aware optimization objectives. However, these methods can be impractical due to significant computational resources, complex industrial tests, and the associated CO2 footprint. Additionally, regular users often fail to fine-tune models because they lack access to model parameters In this paper, we introduce the Inference-Time Rule Eraser (Eraser), a novel method designed to address fairness concerns by removing biased decision-making rules from deployed models during inference without altering model weights. We begin by establishing a theoretical foundation for modifying model outputs to eliminate biased rules through Bayesian analysis. Next, we present a specific implementation of Eraser that involves two stages: (1) distilling the biased rules from the deployed model into an additional patch model, and (2) removing these biased rules from the output of the deployed model during inference. Extensive experiments validate the effectiveness of our approach, showcasing its superior performance in addressing fairness concerns in AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2404_04814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Inference-Time Rule Eraser: Fair Recognition via Distilling and Removing Biased Rules
Zhang, Yi
Lu, Dongyuan
Sang, Jitao
Machine Learning
Artificial Intelligence
Computers and Society
Machine learning models often make predictions based on biased features such as gender, race, and other social attributes, posing significant fairness risks, especially in societal applications, such as hiring, banking, and criminal justice. Traditional approaches to addressing this issue involve retraining or fine-tuning neural networks with fairness-aware optimization objectives. However, these methods can be impractical due to significant computational resources, complex industrial tests, and the associated CO2 footprint. Additionally, regular users often fail to fine-tune models because they lack access to model parameters In this paper, we introduce the Inference-Time Rule Eraser (Eraser), a novel method designed to address fairness concerns by removing biased decision-making rules from deployed models during inference without altering model weights. We begin by establishing a theoretical foundation for modifying model outputs to eliminate biased rules through Bayesian analysis. Next, we present a specific implementation of Eraser that involves two stages: (1) distilling the biased rules from the deployed model into an additional patch model, and (2) removing these biased rules from the output of the deployed model during inference. Extensive experiments validate the effectiveness of our approach, showcasing its superior performance in addressing fairness concerns in AI systems.
title Inference-Time Rule Eraser: Fair Recognition via Distilling and Removing Biased Rules
topic Machine Learning
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2404.04814