Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector Quantization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Luu, Tung M., Nguyen, Thanh, Jin, Tee Joshua Tian, Kim, Sungwoon, Yoo, Chang D.
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909336343674880
author Luu, Tung M.
Nguyen, Thanh
Jin, Tee Joshua Tian
Kim, Sungwoon
Yoo, Chang D.
author_facet Luu, Tung M.
Nguyen, Thanh
Jin, Tee Joshua Tian
Kim, Sungwoon
Yoo, Chang D.
contents Recent studies reveal that well-performing reinforcement learning (RL) agents in training often lack resilience against adversarial perturbations during deployment. This highlights the importance of building a robust agent before deploying it in the real world. Most prior works focus on developing robust training-based procedures to tackle this problem, including enhancing the robustness of the deep neural network component itself or adversarially training the agent on strong attacks. In this work, we instead study an input transformation-based defense for RL. Specifically, we propose using a variant of vector quantization (VQ) as a transformation for input observations, which is then used to reduce the space of adversarial attacks during testing, resulting in the transformed observations being less affected by attacks. Our method is computationally efficient and seamlessly integrates with adversarial training, further enhancing the robustness of RL agents against adversarial attacks. Through extensive experiments in multiple environments, we demonstrate that using VQ as the input transformation effectively defends against adversarial attacks on the agent's observations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03376
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector Quantization
Luu, Tung M.
Nguyen, Thanh
Jin, Tee Joshua Tian
Kim, Sungwoon
Yoo, Chang D.
Machine Learning
Artificial Intelligence
Recent studies reveal that well-performing reinforcement learning (RL) agents in training often lack resilience against adversarial perturbations during deployment. This highlights the importance of building a robust agent before deploying it in the real world. Most prior works focus on developing robust training-based procedures to tackle this problem, including enhancing the robustness of the deep neural network component itself or adversarially training the agent on strong attacks. In this work, we instead study an input transformation-based defense for RL. Specifically, we propose using a variant of vector quantization (VQ) as a transformation for input observations, which is then used to reduce the space of adversarial attacks during testing, resulting in the transformed observations being less affected by attacks. Our method is computationally efficient and seamlessly integrates with adversarial training, further enhancing the robustness of RL agents against adversarial attacks. Through extensive experiments in multiple environments, we demonstrate that using VQ as the input transformation effectively defends against adversarial attacks on the agent's observations.
title Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector Quantization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.03376