Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yamabe, Shojiro, Fukuchi, Kazuto, Sakuma, Jun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910024022884352
author Yamabe, Shojiro
Fukuchi, Kazuto
Sakuma, Jun
author_facet Yamabe, Shojiro
Fukuchi, Kazuto
Sakuma, Jun
contents This study investigates behavior-targeted attacks on reinforcement learning and their countermeasures. Behavior-targeted attacks aim to manipulate the victim's behavior as desired by the adversary through adversarial interventions in state observations. Existing behavior-targeted attacks have some limitations, such as requiring white-box access to the victim's policy. To address this, we propose a novel attack method using imitation learning from adversarial demonstrations, which works under limited access to the victim's policy and is environment-agnostic. In addition, our theoretical analysis proves that the policy's sensitivity to state changes impacts defense performance, particularly in the early stages of the trajectory. Based on this insight, we propose time-discounted regularization, which enhances robustness against attacks while maintaining task performance. To the best of our knowledge, this is the first defense strategy specifically designed for behavior-targeted attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2406_03862
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
Yamabe, Shojiro
Fukuchi, Kazuto
Sakuma, Jun
Machine Learning
Artificial Intelligence
This study investigates behavior-targeted attacks on reinforcement learning and their countermeasures. Behavior-targeted attacks aim to manipulate the victim's behavior as desired by the adversary through adversarial interventions in state observations. Existing behavior-targeted attacks have some limitations, such as requiring white-box access to the victim's policy. To address this, we propose a novel attack method using imitation learning from adversarial demonstrations, which works under limited access to the victim's policy and is environment-agnostic. In addition, our theoretical analysis proves that the policy's sensitivity to state changes impacts defense performance, particularly in the early stages of the trajectory. Based on this insight, we propose time-discounted regularization, which enhances robustness against attacks while maintaining task performance. To the best of our knowledge, this is the first defense strategy specifically designed for behavior-targeted attacks.
title Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.03862