Unlearning Backdoor Attacks through Gradient-Based Model Pruning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dunnett, Kealan, Arablouei, Reza, Miller, Dimity, Dedeoglu, Volkan, Jurdak, Raja
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911869932929024
author Dunnett, Kealan
Arablouei, Reza
Miller, Dimity
Dedeoglu, Volkan
Jurdak, Raja
author_facet Dunnett, Kealan
Arablouei, Reza
Miller, Dimity
Dedeoglu, Volkan
Jurdak, Raja
contents In the era of increasing concerns over cybersecurity threats, defending against backdoor attacks is paramount in ensuring the integrity and reliability of machine learning models. However, many existing approaches require substantial amounts of data for effective mitigation, posing significant challenges in practical deployment. To address this, we propose a novel approach to counter backdoor attacks by treating their mitigation as an unlearning task. We tackle this challenge through a targeted model pruning strategy, leveraging unlearning loss gradients to identify and eliminate backdoor elements within the model. Built on solid theoretical insights, our approach offers simplicity and effectiveness, rendering it well-suited for scenarios with limited data availability. Our methodology includes formulating a suitable unlearning loss and devising a model-pruning technique tailored for convolutional neural networks. Comprehensive evaluations demonstrate the efficacy of our proposed approach compared to state-of-the-art approaches, particularly in realistic data settings.
format Preprint
id arxiv_https___arxiv_org_abs_2405_03918
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unlearning Backdoor Attacks through Gradient-Based Model Pruning
Dunnett, Kealan
Arablouei, Reza
Miller, Dimity
Dedeoglu, Volkan
Jurdak, Raja
Machine Learning
Cryptography and Security
In the era of increasing concerns over cybersecurity threats, defending against backdoor attacks is paramount in ensuring the integrity and reliability of machine learning models. However, many existing approaches require substantial amounts of data for effective mitigation, posing significant challenges in practical deployment. To address this, we propose a novel approach to counter backdoor attacks by treating their mitigation as an unlearning task. We tackle this challenge through a targeted model pruning strategy, leveraging unlearning loss gradients to identify and eliminate backdoor elements within the model. Built on solid theoretical insights, our approach offers simplicity and effectiveness, rendering it well-suited for scenarios with limited data availability. Our methodology includes formulating a suitable unlearning loss and devising a model-pruning technique tailored for convolutional neural networks. Comprehensive evaluations demonstrate the efficacy of our proposed approach compared to state-of-the-art approaches, particularly in realistic data settings.
title Unlearning Backdoor Attacks through Gradient-Based Model Pruning
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2405.03918