How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cheng, Xiaoyuan, Yuan, Wenxuan, Li, Boyang, Xu, Yuanchao, Yang, Yiming, Liang, Hao, Peng, Bei, Loftin, Robert, Sun, Zhuo, Hu, Yukun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913091657138176
author Cheng, Xiaoyuan
Yuan, Wenxuan
Li, Boyang
Xu, Yuanchao
Yang, Yiming
Liang, Hao
Peng, Bei
Loftin, Robert
Sun, Zhuo
Hu, Yukun
author_facet Cheng, Xiaoyuan
Yuan, Wenxuan
Li, Boyang
Xu, Yuanchao
Yang, Yiming
Liang, Hao
Peng, Bei
Loftin, Robert
Sun, Zhuo
Hu, Yukun
contents Diffusion policy sampling enables reinforcement learning (RL) to represent multimodal action distributions beyond suboptimal unimodal Gaussian policies. However, existing diffusion-based RL methods primarily focus on offline settings for reward maximization, with limited consideration of safety in online settings. To address this gap, we propose Augmented Lagrangian-Guided Diffusion (ALGD), a novel algorithm for off-policy safe RL. By revisiting optimization theory and energy-based model, we show that the instability of primal-dual methods arises from the non-convex Lagrangian landscape. In diffusion-based safe RL, the Lagrangian can be interpreted as an energy function guiding the denoising dynamics. Counterintuitively, direct usage destabilizes both policy generation and training. ALGD resolves this issue by introducing an augmented Lagrangian that locally convexifies the energy landscape, yielding a stabilized policy generation and training process without altering the distribution of the optimal policy. Theoretical analysis and extensive experiments demonstrate that ALGD is both theoretically grounded and empirically effective, achieving strong and stable performance across diverse environments.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02924
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?
Cheng, Xiaoyuan
Yuan, Wenxuan
Li, Boyang
Xu, Yuanchao
Yang, Yiming
Liang, Hao
Peng, Bei
Loftin, Robert
Sun, Zhuo
Hu, Yukun
Machine Learning
Systems and Control
Diffusion policy sampling enables reinforcement learning (RL) to represent multimodal action distributions beyond suboptimal unimodal Gaussian policies. However, existing diffusion-based RL methods primarily focus on offline settings for reward maximization, with limited consideration of safety in online settings. To address this gap, we propose Augmented Lagrangian-Guided Diffusion (ALGD), a novel algorithm for off-policy safe RL. By revisiting optimization theory and energy-based model, we show that the instability of primal-dual methods arises from the non-convex Lagrangian landscape. In diffusion-based safe RL, the Lagrangian can be interpreted as an energy function guiding the denoising dynamics. Counterintuitively, direct usage destabilizes both policy generation and training. ALGD resolves this issue by introducing an augmented Lagrangian that locally convexifies the energy landscape, yielding a stabilized policy generation and training process without altering the distribution of the optimal policy. Theoretical analysis and extensive experiments demonstrate that ALGD is both theoretically grounded and empirically effective, achieving strong and stable performance across diverse environments.
title How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?
topic Machine Learning
Systems and Control
url https://arxiv.org/abs/2602.02924