When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Flynn, Donald, Goldhirsh, Hadas Yaron, Keating, Jonathan P., Seroussi, Inbar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914588135522304
author Flynn, Donald
Goldhirsh, Hadas Yaron
Keating, Jonathan P.
Seroussi, Inbar
author_facet Flynn, Donald
Goldhirsh, Hadas Yaron
Keating, Jonathan P.
Seroussi, Inbar
contents Backdoor poisoning attacks behave counter-intuitively in high dimensions: stronger training triggers can help the defender. We study regularised generalised linear models on Gaussian-mixture data in the proportional regime ($p/n \to κ$), varying the training trigger strength $α$ against a fixed test trigger. Three phenomena emerge: (i) clean test accuracy increases with $α$; (ii) attack success peaks at a finite $α$ and then declines; and (iii) the most damaging trigger direction is the minimum eigenvector of the data covariance. We prove all three results in closed form for the squared loss, and extend (i) and (ii) to general convex GLM losses via a Gaussian-proxy fixed-point system. We identify a finite-sample noise floor proportional to $κ$ as the mechanism behind (i), invisible to classical $n \gg p$ analysis. Experiments on CIFAR-10 and Gaussian surrogates match the theory closely; ResNet-18 experiments show the same phenomena beyond the convex setting.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22481
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks
Flynn, Donald
Goldhirsh, Hadas Yaron
Keating, Jonathan P.
Seroussi, Inbar
Machine Learning
Statistics Theory
Backdoor poisoning attacks behave counter-intuitively in high dimensions: stronger training triggers can help the defender. We study regularised generalised linear models on Gaussian-mixture data in the proportional regime ($p/n \to κ$), varying the training trigger strength $α$ against a fixed test trigger. Three phenomena emerge: (i) clean test accuracy increases with $α$; (ii) attack success peaks at a finite $α$ and then declines; and (iii) the most damaging trigger direction is the minimum eigenvector of the data covariance. We prove all three results in closed form for the squared loss, and extend (i) and (ii) to general convex GLM losses via a Gaussian-proxy fixed-point system. We identify a finite-sample noise floor proportional to $κ$ as the mechanism behind (i), invisible to classical $n \gg p$ analysis. Experiments on CIFAR-10 and Gaussian surrogates match the theory closely; ResNet-18 experiments show the same phenomena beyond the convex setting.
title When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks
topic Machine Learning
Statistics Theory
url https://arxiv.org/abs/2605.22481