Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Guo, Junyu, Zheng, Zhi, Ying, Donghao, Jin, Ming, Gu, Shangding, Spanos, Costas, Lavaei, Javad
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911138615132160
author Guo, Junyu
Zheng, Zhi
Ying, Donghao
Jin, Ming
Gu, Shangding
Spanos, Costas
Lavaei, Javad
author_facet Guo, Junyu
Zheng, Zhi
Ying, Donghao
Jin, Ming
Gu, Shangding
Spanos, Costas
Lavaei, Javad
contents Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent has only a fixed dataset -- common in realistic tasks to prevent unsafe exploration. To address this, we propose Diffusion-Regularized Constrained Offline Reinforcement Learning (DRCORL), which first uses a diffusion model to capture the behavioral policy from offline data and then extracts a simplified policy to enable efficient inference. We further apply gradient manipulation for safety adaptation, balancing the reward objective and constraint satisfaction. This approach leverages high-quality offline data while incorporating safety requirements. Empirical results show that DRCORL achieves reliable safety performance, fast inference, and strong reward outcomes across robot learning tasks. Compared to existing safe offline RL methods, it consistently meets cost limits and performs well with the same hyperparameters, indicating practical applicability in real-world scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12391
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
Guo, Junyu
Zheng, Zhi
Ying, Donghao
Jin, Ming
Gu, Shangding
Spanos, Costas
Lavaei, Javad
Machine Learning
Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent has only a fixed dataset -- common in realistic tasks to prevent unsafe exploration. To address this, we propose Diffusion-Regularized Constrained Offline Reinforcement Learning (DRCORL), which first uses a diffusion model to capture the behavioral policy from offline data and then extracts a simplified policy to enable efficient inference. We further apply gradient manipulation for safety adaptation, balancing the reward objective and constraint satisfaction. This approach leverages high-quality offline data while incorporating safety requirements. Empirical results show that DRCORL achieves reliable safety performance, fast inference, and strong reward outcomes across robot learning tasks. Compared to existing safe offline RL methods, it consistently meets cost limits and performs well with the same hyperparameters, indicating practical applicability in real-world scenarios.
title Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
topic Machine Learning
url https://arxiv.org/abs/2502.12391