Policy Bifurcation in Safe Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Wenjun, Lyu, Yao, Li, Jie, Yang, Yujie, Li, Shengbo Eben, Duan, Jingliang, Zhan, Xianyuan, Liu, Jingjing, Zhang, Yaqin, Li, Keqiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feasible Policy Iteration for Safe Reinforcement Learning
by: Yang, Yujie, et al.
Published: (2023)
by: Yang, Yujie, et al.
Published: (2023)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
by: Zheng, Yinan, et al.
Published: (2024)
by: Zheng, Yinan, et al.
Published: (2024)
Conformal Symplectic Optimization for Stable Reinforcement Learning
by: Lyu, Yao, et al.
Published: (2024)
by: Lyu, Yao, et al.
Published: (2024)
Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
by: Zhang, Feihong, et al.
Published: (2025)
by: Zhang, Feihong, et al.
Published: (2025)
Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model
by: Guan, Yang, et al.
Published: (2021)
by: Guan, Yang, et al.
Published: (2021)
Predictive Lagrangian Optimization for Constrained Reinforcement Learning
by: Zhang, Tianqi, et al.
Published: (2025)
by: Zhang, Tianqi, et al.
Published: (2025)
On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration
by: Yang, Yujie, et al.
Published: (2026)
by: Yang, Yujie, et al.
Published: (2026)
Jump-Start Reinforcement Learning with Self-Evolving Priors for Extreme Monopedal Locomotion
by: Zheng, Ziang, et al.
Published: (2025)
by: Zheng, Ziang, et al.
Published: (2025)
Distributional Soft Actor-Critic with Diffusion Policy
by: Liu, Tong, et al.
Published: (2025)
by: Liu, Tong, et al.
Published: (2025)
Distributional Soft Actor-Critic with Three Refinements
by: Duan, Jingliang, et al.
Published: (2023)
by: Duan, Jingliang, et al.
Published: (2023)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Real-Time Generative Policy via Langevin-Guided Flow Matching for Autonomous Driving
by: Zhu, Tianze, et al.
Published: (2026)
by: Zhu, Tianze, et al.
Published: (2026)
Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model
by: Gao, Zeyu, et al.
Published: (2022)
by: Gao, Zeyu, et al.
Published: (2022)
Rocket Landing Control with Random Annealing Jump Start Reinforcement Learning
by: Jiang, Yuxuan, et al.
Published: (2024)
by: Jiang, Yuxuan, et al.
Published: (2024)
Query-Policy Misalignment in Preference-Based Reinforcement Learning
by: Hu, Xiao, et al.
Published: (2023)
by: Hu, Xiao, et al.
Published: (2023)
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
by: Wang, Likun, et al.
Published: (2025)
by: Wang, Likun, et al.
Published: (2025)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
by: Liu, Shiqi, et al.
Published: (2026)
by: Liu, Shiqi, et al.
Published: (2026)
Diffusion Actor-Critic with Entropy Regulator
by: Wang, Yinuo, et al.
Published: (2024)
by: Wang, Yinuo, et al.
Published: (2024)
A Spectral Revisit of the Distributional Bellman Operator under the Cramér Metric
by: Wang, Keru, et al.
Published: (2026)
by: Wang, Keru, et al.
Published: (2026)
Enhanced DACER Algorithm with High Diffusion Efficiency
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
Diffusion-Based Planning for Autonomous Driving with Flexible Guidance
by: Zheng, Yinan, et al.
Published: (2025)
by: Zheng, Yinan, et al.
Published: (2025)
Bidirectional-Reachable Hierarchical Reinforcement Learning with Mutually Responsive Policies
by: Luo, Yu, et al.
Published: (2024)
by: Luo, Yu, et al.
Published: (2024)
The Feasibility Theory of Constrained Reinforcement Learning: A Tutorial Study
by: Yang, Yujie, et al.
Published: (2024)
by: Yang, Yujie, et al.
Published: (2024)
Hierarchical End-to-End Autonomous Driving: Integrating BEV Perception with Deep Reinforcement Learning
by: Lu, Siyi, et al.
Published: (2024)
by: Lu, Siyi, et al.
Published: (2024)
Hierarchical Feature-level Reverse Propagation for Post-Training Neural Networks
by: Ding, Ni, et al.
Published: (2025)
by: Ding, Ni, et al.
Published: (2025)
Zeroth-Order Actor-Critic: An Evolutionary Framework for Sequential Decision Problems
by: Lei, Yuheng, et al.
Published: (2022)
by: Lei, Yuheng, et al.
Published: (2022)
On the Stability of Datatic Control Systems
by: Yang, Yujie, et al.
Published: (2024)
by: Yang, Yujie, et al.
Published: (2024)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
by: Zhan, Guojian, et al.
Published: (2026)
by: Zhan, Guojian, et al.
Published: (2026)
On the Optimization Landscape of Observer-based Dynamic Linear Quadratic Control
by: Duan, Jingliang, et al.
Published: (2026)
by: Duan, Jingliang, et al.
Published: (2026)
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
DRARL: Disengagement-Reason-Augmented Reinforcement Learning for Efficient Improvement of Autonomous Driving Policy
by: Zhou, Weitao, et al.
Published: (2025)
by: Zhou, Weitao, et al.
Published: (2025)
Diffusion-DICE: In-Sample Diffusion Guidance for Offline Reinforcement Learning
by: Mao, Liyuan, et al.
Published: (2024)
by: Mao, Liyuan, et al.
Published: (2024)
Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling
by: Zhou, Xubin, et al.
Published: (2026)
by: Zhou, Xubin, et al.
Published: (2026)
Canonical Form of Datatic Description in Control Systems
by: Zhan, Guojian, et al.
Published: (2024)
by: Zhan, Guojian, et al.
Published: (2024)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
by: Yao, Yihang, et al.
Published: (2023)
by: Yao, Yihang, et al.
Published: (2023)
Controllability Test for Nonlinear Datatic Systems
by: Yang, Yujie, et al.
Published: (2024)
by: Yang, Yujie, et al.
Published: (2024)
Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies
by: Li, Jin, et al.
Published: (2026)
by: Li, Jin, et al.
Published: (2026)
Natural Gradient Gaussian Approximation Filter on Lie Groups for Robot State Estimation
by: Zhang, Tianyi, et al.
Published: (2026)
by: Zhang, Tianyi, et al.
Published: (2026)
Similar Items
-
Feasible Policy Iteration for Safe Reinforcement Learning
by: Yang, Yujie, et al.
Published: (2023) -
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
by: Zheng, Yinan, et al.
Published: (2024) -
Conformal Symplectic Optimization for Stable Reinforcement Learning
by: Lyu, Yao, et al.
Published: (2024) -
Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
by: Zhang, Jiaming, et al.
Published: (2025) -
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning
by: Zhang, Jiaming, et al.
Published: (2026)