On the existence of consistent adversarial attacks in high-dimensional linear classification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vilucchio, Matteo, Zdeborová, Lenka, Loureiro, Bruno
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909649180033024
author Vilucchio, Matteo
Zdeborová, Lenka
Loureiro, Bruno
author_facet Vilucchio, Matteo
Zdeborová, Lenka
Loureiro, Bruno
contents What fundamentally distinguishes an adversarial attack from a misclassification due to limited model expressivity or finite data? In this work, we investigate this question in the setting of high-dimensional binary classification, where statistical effects due to limited data availability play a central role. We introduce a new error metric that precisely capture this distinction, quantifying model vulnerability to consistent adversarial attacks -- perturbations that preserve the ground-truth labels. Our main technical contribution is an exact and rigorous asymptotic characterization of these metrics in both well-specified models and latent space models, revealing different vulnerability patterns compared to standard robust error measures. The theoretical results demonstrate that as models become more overparameterized, their vulnerability to label-preserving perturbations grows, offering theoretical insight into the mechanisms underlying model sensitivity to adversarial attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_12454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the existence of consistent adversarial attacks in high-dimensional linear classification
Vilucchio, Matteo
Zdeborová, Lenka
Loureiro, Bruno
Machine Learning
Disordered Systems and Neural Networks
Cryptography and Security
What fundamentally distinguishes an adversarial attack from a misclassification due to limited model expressivity or finite data? In this work, we investigate this question in the setting of high-dimensional binary classification, where statistical effects due to limited data availability play a central role. We introduce a new error metric that precisely capture this distinction, quantifying model vulnerability to consistent adversarial attacks -- perturbations that preserve the ground-truth labels. Our main technical contribution is an exact and rigorous asymptotic characterization of these metrics in both well-specified models and latent space models, revealing different vulnerability patterns compared to standard robust error measures. The theoretical results demonstrate that as models become more overparameterized, their vulnerability to label-preserving perturbations grows, offering theoretical insight into the mechanisms underlying model sensitivity to adversarial attacks.
title On the existence of consistent adversarial attacks in high-dimensional linear classification
topic Machine Learning
Disordered Systems and Neural Networks
Cryptography and Security
url https://arxiv.org/abs/2506.12454