The Adversarial Conditioning Paradox: Why Attacked Inputs Are More Stable, Not Less

Fuente: Zenodo
Saved in:
Bibliographic Details
Main Authors: Sapenov, Khazretgali, Sapenov, Aidos
Format: Recurso digital
Language:English
Published: Zenodo 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866901180815245312
author Sapenov, Khazretgali
Sapenov, Aidos
author_facet Sapenov, Khazretgali
Sapenov, Aidos
contents <p>Adversarial attacks on NLP systems are designed to find inputs that fool models while minimizing perceptible changes, making them difficult to detect using similarity-based methods. We investigate whether Jacobian conditioning analysis can provide an orthogonal detection signal. Surprisingly, we find that adversarial inputs exhibit systematically lower condition numbers at early transformer layers—the opposite of our initial hypothesis that attacks exploit unstable, ill-conditioned regions. This “adversarial conditioning paradox” replicates across multiple attack types: TextFooler (AUC = 0.72, p = 0.001), DeepWordBug (AUC = 0.75, p = 0.001), and directionally for PWWS (AUC = 0.59, p = 0.29). The effect holds for both word-level and character-level perturbations, while embedding cosine distance fails completely (AUC ≈ 0.25). We propose that adversarial attacks succeed by finding wellconditioned directions that cross decision boundaries—smooth paths to misclassification rather than chaotic exploitation of instability. Our findings open new directions for adversarial detection using internal geometric properties invisible to embedding-based methods.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_17818479
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle The Adversarial Conditioning Paradox: Why Attacked Inputs Are More Stable, Not Less
Sapenov, Khazretgali
Sapenov, Aidos
adversarial detection, Jacobian conditioning, transformer robustness, NLP security, condition number analysis
<p>Adversarial attacks on NLP systems are designed to find inputs that fool models while minimizing perceptible changes, making them difficult to detect using similarity-based methods. We investigate whether Jacobian conditioning analysis can provide an orthogonal detection signal. Surprisingly, we find that adversarial inputs exhibit systematically lower condition numbers at early transformer layers—the opposite of our initial hypothesis that attacks exploit unstable, ill-conditioned regions. This “adversarial conditioning paradox” replicates across multiple attack types: TextFooler (AUC = 0.72, p = 0.001), DeepWordBug (AUC = 0.75, p = 0.001), and directionally for PWWS (AUC = 0.59, p = 0.29). The effect holds for both word-level and character-level perturbations, while embedding cosine distance fails completely (AUC ≈ 0.25). We propose that adversarial attacks succeed by finding wellconditioned directions that cross decision boundaries—smooth paths to misclassification rather than chaotic exploitation of instability. Our findings open new directions for adversarial detection using internal geometric properties invisible to embedding-based methods.</p>
title The Adversarial Conditioning Paradox: Why Attacked Inputs Are More Stable, Not Less
topic adversarial detection, Jacobian conditioning, transformer robustness, NLP security, condition number analysis
url https://doi.org/10.5281/zenodo.17818479