Defending Quantum Classifiers against Adversarial Perturbations through Quantum Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Andrews, Emma, Sanjaya, Sahan, Mishra, Prabhat
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909005233782784
author Andrews, Emma
Sanjaya, Sahan
Mishra, Prabhat
author_facet Andrews, Emma
Sanjaya, Sahan
Mishra, Prabhat
contents Machine learning models can learn from data samples to carry out various tasks efficiently. When data samples are adversarially manipulated, such as by insertion of carefully crafted noise, it can cause the model to make mistakes. Quantum machine learning models are also vulnerable to such adversarial attacks, especially in image classification using variational quantum classifiers. While there are promising defenses against these adversarial perturbations, such as training with adversarial samples, they face practical limitations. For example, they are not applicable in scenarios where training with adversarial samples is either not possible or can overfit the models on one type of attack. In this paper, we propose an adversarial training-free defense framework that utilizes a quantum autoencoder to purify the adversarial samples through reconstruction. Moreover, our defense framework provides a confidence metric to identify potentially adversarial samples that cannot be purified the quantum autoencoder. Extensive evaluation demonstrates that our defense framework can significantly outperform state-of-the-art in prediction accuracy (up to 68%) under adversarial attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_28176
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Defending Quantum Classifiers against Adversarial Perturbations through Quantum Autoencoders
Andrews, Emma
Sanjaya, Sahan
Mishra, Prabhat
Quantum Physics
Machine Learning
Machine learning models can learn from data samples to carry out various tasks efficiently. When data samples are adversarially manipulated, such as by insertion of carefully crafted noise, it can cause the model to make mistakes. Quantum machine learning models are also vulnerable to such adversarial attacks, especially in image classification using variational quantum classifiers. While there are promising defenses against these adversarial perturbations, such as training with adversarial samples, they face practical limitations. For example, they are not applicable in scenarios where training with adversarial samples is either not possible or can overfit the models on one type of attack. In this paper, we propose an adversarial training-free defense framework that utilizes a quantum autoencoder to purify the adversarial samples through reconstruction. Moreover, our defense framework provides a confidence metric to identify potentially adversarial samples that cannot be purified the quantum autoencoder. Extensive evaluation demonstrates that our defense framework can significantly outperform state-of-the-art in prediction accuracy (up to 68%) under adversarial attacks.
title Defending Quantum Classifiers against Adversarial Perturbations through Quantum Autoencoders
topic Quantum Physics
Machine Learning
url https://arxiv.org/abs/2604.28176