DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Popovic, Dorde, Sadeghi, Amin, Yu, Ting, Chawla, Sanjay, Khalil, Issa
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910896240984064
author Popovic, Dorde
Sadeghi, Amin
Yu, Ting
Chawla, Sanjay
Khalil, Issa
author_facet Popovic, Dorde
Sadeghi, Amin
Yu, Ting
Chawla, Sanjay
Khalil, Issa
contents Backdoor attacks are among the most effective, practical, and stealthy attacks in deep learning. In this paper, we consider a practical scenario where a developer obtains a deep model from a third party and uses it as part of a safety-critical system. The developer wants to inspect the model for potential backdoors prior to system deployment. We find that most existing detection techniques make assumptions that are not applicable to this scenario. In this paper, we present a novel framework for detecting backdoors under realistic restrictions. We generate candidate triggers by deductively searching over the space of possible triggers. We construct and optimize a smoothed version of Attack Success Rate as our search objective. Starting from a broad class of template attacks and just using the forward pass of a deep model, we reverse engineer the backdoor attack. We conduct extensive evaluation on a wide range of attacks, models, and datasets, with our technique performing almost perfectly across these settings.
format Preprint
id arxiv_https___arxiv_org_abs_2503_21305
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
Popovic, Dorde
Sadeghi, Amin
Yu, Ting
Chawla, Sanjay
Khalil, Issa
Cryptography and Security
Artificial Intelligence
Backdoor attacks are among the most effective, practical, and stealthy attacks in deep learning. In this paper, we consider a practical scenario where a developer obtains a deep model from a third party and uses it as part of a safety-critical system. The developer wants to inspect the model for potential backdoors prior to system deployment. We find that most existing detection techniques make assumptions that are not applicable to this scenario. In this paper, we present a novel framework for detecting backdoors under realistic restrictions. We generate candidate triggers by deductively searching over the space of possible triggers. We construct and optimize a smoothed version of Attack Success Rate as our search objective. Starting from a broad class of template attacks and just using the forward pass of a deep model, we reverse engineer the backdoor attack. We conduct extensive evaluation on a wide range of attacks, models, and datasets, with our technique performing almost perfectly across these settings.
title DeBackdoor: A Deductive Framework for Detecting Backdoor Attacks on Deep Models with Limited Data
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2503.21305