Vulnerability Analysis of Safe Reinforcement Learning via Inverse Constrained Reinforcement Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fan, Jialiang, Jiang, Shixiong, Liu, Mengyu, Kong, Fanxin
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908839859716096
author Fan, Jialiang
Jiang, Shixiong
Liu, Mengyu
Kong, Fanxin
author_facet Fan, Jialiang
Jiang, Shixiong
Liu, Mengyu
Kong, Fanxin
contents Safe reinforcement learning (Safe RL) aims to ensure policy performance while satisfying safety constraints. However, most existing Safe RL methods assume benign environments, making them vulnerable to adversarial perturbations commonly encountered in real-world settings. In addition, existing gradient-based adversarial attacks typically require access to the policy's gradient information, which is often impractical in real-world scenarios. To address these challenges, we propose an adversarial attack framework to reveal vulnerabilities of Safe RL policies. Using expert demonstrations and black-box environment interaction, our framework learns a constraint model and a surrogate (learner) policy, enabling gradient-based attack optimization without requiring the victim policy's internal gradients or the ground-truth safety constraints. We further provide theoretical analysis establishing feasibility and deriving perturbation bounds. Experiments on multiple Safe RL benchmarks demonstrate the effectiveness of our approach under limited privileged access.
format Preprint
id arxiv_https___arxiv_org_abs_2602_16543
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Vulnerability Analysis of Safe Reinforcement Learning via Inverse Constrained Reinforcement Learning
Fan, Jialiang
Jiang, Shixiong
Liu, Mengyu
Kong, Fanxin
Machine Learning
Safe reinforcement learning (Safe RL) aims to ensure policy performance while satisfying safety constraints. However, most existing Safe RL methods assume benign environments, making them vulnerable to adversarial perturbations commonly encountered in real-world settings. In addition, existing gradient-based adversarial attacks typically require access to the policy's gradient information, which is often impractical in real-world scenarios. To address these challenges, we propose an adversarial attack framework to reveal vulnerabilities of Safe RL policies. Using expert demonstrations and black-box environment interaction, our framework learns a constraint model and a surrogate (learner) policy, enabling gradient-based attack optimization without requiring the victim policy's internal gradients or the ground-truth safety constraints. We further provide theoretical analysis establishing feasibility and deriving perturbation bounds. Experiments on multiple Safe RL benchmarks demonstrate the effectiveness of our approach under limited privileged access.
title Vulnerability Analysis of Safe Reinforcement Learning via Inverse Constrained Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2602.16543