Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mehrotra, Ayushi, Peng, Derek, Bhusal, Dipkamal, Rastogi, Nidhi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914076192407552
author Mehrotra, Ayushi
Peng, Derek
Bhusal, Dipkamal
Rastogi, Nidhi
author_facet Mehrotra, Ayushi
Peng, Derek
Bhusal, Dipkamal
Rastogi, Nidhi
contents Adversarial patch attacks pose a practical threat to deep learning models by forcing targeted misclassifications through localized perturbations, often realized in the physical world. Existing defenses typically assume prior knowledge of patch size or location, limiting their applicability. In this work, we propose a patch-agnostic defense that leverages concept-based explanations to identify and suppress the most influential concept activation vectors, thereby neutralizing patch effects without explicit detection. Evaluated on Imagenette with a ResNet-50, our method achieves higher robust and clean accuracy than the state-of-the-art PatchCleanser, while maintaining strong performance across varying patch sizes and locations. Our results highlight the promise of combining interpretability with robustness and suggest concept-driven defenses as a scalable strategy for securing machine learning models against adversarial patch attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2510_04245
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
Mehrotra, Ayushi
Peng, Derek
Bhusal, Dipkamal
Rastogi, Nidhi
Computer Vision and Pattern Recognition
Artificial Intelligence
Adversarial patch attacks pose a practical threat to deep learning models by forcing targeted misclassifications through localized perturbations, often realized in the physical world. Existing defenses typically assume prior knowledge of patch size or location, limiting their applicability. In this work, we propose a patch-agnostic defense that leverages concept-based explanations to identify and suppress the most influential concept activation vectors, thereby neutralizing patch effects without explicit detection. Evaluated on Imagenette with a ResNet-50, our method achieves higher robust and clean accuracy than the state-of-the-art PatchCleanser, while maintaining strong performance across varying patch sizes and locations. Our results highlight the promise of combining interpretability with robustness and suggest concept-driven defenses as a scalable strategy for securing machine learning models against adversarial patch attacks.
title Concept-Based Masking: A Patch-Agnostic Defense Against Adversarial Patch Attacks
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.04245