Frequency maps reveal the correlation between Adversarial Attacks and Implicit Bias

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Basile, Lorenzo, Karantzas, Nikos, d'Onofrio, Alberto, Manzoni, Luca, Bortolussi, Luca, Rodriguez, Alex, Anselmi, Fabio
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916677098143744
author Basile, Lorenzo
Karantzas, Nikos
d'Onofrio, Alberto
Manzoni, Luca
Bortolussi, Luca
Rodriguez, Alex
Anselmi, Fabio
author_facet Basile, Lorenzo
Karantzas, Nikos
d'Onofrio, Alberto
Manzoni, Luca
Bortolussi, Luca
Rodriguez, Alex
Anselmi, Fabio
contents Despite their impressive performance in classification tasks, neural networks are known to be vulnerable to adversarial attacks, subtle perturbations of the input data designed to deceive the model. In this work, we investigate the correlation between these perturbations and the implicit bias of neural networks trained with gradient-based algorithms. To this end, we analyse a representation of the network's implicit bias through the lens of the Fourier transform. Specifically, we identify unique fingerprints of implicit bias and adversarial attacks by calculating the minimal, essential frequencies needed for accurate classification of each image, as well as the frequencies that drive misclassification in its adversarially perturbed counterpart. This approach enables us to uncover and analyse the correlation between these essential frequencies, providing a precise map of how the network's biases align or contrast with the frequency components exploited by adversarial attacks. To this end, among other methods, we use a newly introduced technique capable of detecting nonlinear correlations between high-dimensional datasets. Our results provide empirical evidence that the network bias in Fourier space and the target frequencies of adversarial attacks are highly correlated and suggest new potential strategies for adversarial defence.
format Preprint
id arxiv_https___arxiv_org_abs_2305_15203
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Frequency maps reveal the correlation between Adversarial Attacks and Implicit Bias
Basile, Lorenzo
Karantzas, Nikos
d'Onofrio, Alberto
Manzoni, Luca
Bortolussi, Luca
Rodriguez, Alex
Anselmi, Fabio
Machine Learning
Artificial Intelligence
Cryptography and Security
Despite their impressive performance in classification tasks, neural networks are known to be vulnerable to adversarial attacks, subtle perturbations of the input data designed to deceive the model. In this work, we investigate the correlation between these perturbations and the implicit bias of neural networks trained with gradient-based algorithms. To this end, we analyse a representation of the network's implicit bias through the lens of the Fourier transform. Specifically, we identify unique fingerprints of implicit bias and adversarial attacks by calculating the minimal, essential frequencies needed for accurate classification of each image, as well as the frequencies that drive misclassification in its adversarially perturbed counterpart. This approach enables us to uncover and analyse the correlation between these essential frequencies, providing a precise map of how the network's biases align or contrast with the frequency components exploited by adversarial attacks. To this end, among other methods, we use a newly introduced technique capable of detecting nonlinear correlations between high-dimensional datasets. Our results provide empirical evidence that the network bias in Fourier space and the target frequencies of adversarial attacks are highly correlated and suggest new potential strategies for adversarial defence.
title Frequency maps reveal the correlation between Adversarial Attacks and Implicit Bias
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2305.15203