Pre-trained Multiple Latent Variable Generative Models are good defenders against Adversarial Attacks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Serez, Dario, Cristani, Marco, Del Bue, Alessio, Murino, Vittorio, Morerio, Pietro
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910727359430656
author Serez, Dario
Cristani, Marco
Del Bue, Alessio
Murino, Vittorio
Morerio, Pietro
author_facet Serez, Dario
Cristani, Marco
Del Bue, Alessio
Murino, Vittorio
Morerio, Pietro
contents Attackers can deliberately perturb classifiers' input with subtle noise, altering final predictions. Among proposed countermeasures, adversarial purification employs generative networks to preprocess input images, filtering out adversarial noise. In this study, we propose specific generators, defined Multiple Latent Variable Generative Models (MLVGMs), for adversarial purification. These models possess multiple latent variables that naturally disentangle coarse from fine features. Taking advantage of these properties, we autoencode images to maintain class-relevant information, while discarding and re-sampling any detail, including adversarial noise. The procedure is completely training-free, exploring the generalization abilities of pre-trained MLVGMs on the adversarial purification downstream task. Despite the lack of large models, trained on billions of samples, we show that smaller MLVGMs are already competitive with traditional methods, and can be used as foundation models. Official code released at https://github.com/SerezD/gen_adversarial.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03453
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pre-trained Multiple Latent Variable Generative Models are good defenders against Adversarial Attacks
Serez, Dario
Cristani, Marco
Del Bue, Alessio
Murino, Vittorio
Morerio, Pietro
Computer Vision and Pattern Recognition
Attackers can deliberately perturb classifiers' input with subtle noise, altering final predictions. Among proposed countermeasures, adversarial purification employs generative networks to preprocess input images, filtering out adversarial noise. In this study, we propose specific generators, defined Multiple Latent Variable Generative Models (MLVGMs), for adversarial purification. These models possess multiple latent variables that naturally disentangle coarse from fine features. Taking advantage of these properties, we autoencode images to maintain class-relevant information, while discarding and re-sampling any detail, including adversarial noise. The procedure is completely training-free, exploring the generalization abilities of pre-trained MLVGMs on the adversarial purification downstream task. Despite the lack of large models, trained on billions of samples, we show that smaller MLVGMs are already competitive with traditional methods, and can be used as foundation models. Official code released at https://github.com/SerezD/gen_adversarial.
title Pre-trained Multiple Latent Variable Generative Models are good defenders against Adversarial Attacks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.03453