Saved in:
Bibliographic Details
Main Authors: Sun, Xiaolin, Liu, Feidi, Ding, Zhengming, Zheng, ZiZhan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.07701
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912701130735616
author Sun, Xiaolin
Liu, Feidi
Ding, Zhengming
Zheng, ZiZhan
author_facet Sun, Xiaolin
Liu, Feidi
Ding, Zhengming
Zheng, ZiZhan
contents Reinforcement learning (RL) systems, while achieving remarkable success across various domains, are vulnerable to adversarial attacks. This is especially a concern in vision-based environments where minor manipulations of high-dimensional image inputs can easily mislead the agent's behavior. To this end, various defenses have been proposed recently, with state-of-the-art approaches achieving robust performance even under large state perturbations. However, after closer investigation, we found that the effectiveness of the current defenses is due to a fundamental weakness of the existing $l_p$ norm-constrained attacks, which can barely alter the semantics of image input even under a relatively large perturbation budget. In this work, we propose SHIFT, a novel policy-agnostic diffusion-based state perturbation attack to go beyond this limitation. Our attack is able to generate perturbed states that are semantically different from the true states while remaining realistic and history-aligned to avoid detection. Evaluations show that our attack effectively breaks existing defenses, including the most sophisticated ones, significantly outperforming existing attacks while being more perceptually stealthy. The results highlight the vulnerability of RL agents to semantics-aware adversarial perturbations, indicating the importance of developing more robust policies.
format Preprint
id arxiv_https___arxiv_org_abs_2511_07701
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
Sun, Xiaolin
Liu, Feidi
Ding, Zhengming
Zheng, ZiZhan
Machine Learning
Artificial Intelligence
Reinforcement learning (RL) systems, while achieving remarkable success across various domains, are vulnerable to adversarial attacks. This is especially a concern in vision-based environments where minor manipulations of high-dimensional image inputs can easily mislead the agent's behavior. To this end, various defenses have been proposed recently, with state-of-the-art approaches achieving robust performance even under large state perturbations. However, after closer investigation, we found that the effectiveness of the current defenses is due to a fundamental weakness of the existing $l_p$ norm-constrained attacks, which can barely alter the semantics of image input even under a relatively large perturbation budget. In this work, we propose SHIFT, a novel policy-agnostic diffusion-based state perturbation attack to go beyond this limitation. Our attack is able to generate perturbed states that are semantically different from the true states while remaining realistic and history-aligned to avoid detection. Evaluations show that our attack effectively breaks existing defenses, including the most sophisticated ones, significantly outperforming existing attacks while being more perceptually stealthy. The results highlight the vulnerability of RL agents to semantics-aware adversarial perturbations, indicating the importance of developing more robust policies.
title Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2511.07701