Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Jiaming, Yang, Yujie, Wang, Haoning, Zhang, Liping, Li, Shengbo Eben |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026)
Feasible Policy Iteration for Safe Reinforcement Learning
von: Yang, Yujie, et al.
Veröffentlicht: (2023)
von: Yang, Yujie, et al.
Veröffentlicht: (2023)
Policy Bifurcation in Safe Reinforcement Learning
von: Zou, Wenjun, et al.
Veröffentlicht: (2024)
von: Zou, Wenjun, et al.
Veröffentlicht: (2024)
On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration
von: Yang, Yujie, et al.
Veröffentlicht: (2026)
von: Yang, Yujie, et al.
Veröffentlicht: (2026)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)
Predictive Lagrangian Optimization for Constrained Reinforcement Learning
von: Zhang, Tianqi, et al.
Veröffentlicht: (2025)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2025)
Conformal Symplectic Optimization for Stable Reinforcement Learning
von: Lyu, Yao, et al.
Veröffentlicht: (2024)
von: Lyu, Yao, et al.
Veröffentlicht: (2024)
Rocket Landing Control with Random Annealing Jump Start Reinforcement Learning
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2024)
Jump-Start Reinforcement Learning with Self-Evolving Priors for Extreme Monopedal Locomotion
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
von: Zheng, Ziang, et al.
Veröffentlicht: (2025)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
von: Zhang, Feihong, et al.
Veröffentlicht: (2025)
Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model
von: Guan, Yang, et al.
Veröffentlicht: (2021)
von: Guan, Yang, et al.
Veröffentlicht: (2021)
Extreme Value Policy Optimization for Safe Reinforcement Learning
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
von: Gao, Shiqing, et al.
Veröffentlicht: (2026)
SafeDreamer: Safe Reinforcement Learning with World Models
von: Huang, Weidong, et al.
Veröffentlicht: (2023)
von: Huang, Weidong, et al.
Veröffentlicht: (2023)
Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement Learning
von: Yao, Yihang, et al.
Veröffentlicht: (2023)
von: Yao, Yihang, et al.
Veröffentlicht: (2023)
Real-Time Generative Policy via Langevin-Guided Flow Matching for Autonomous Driving
von: Zhu, Tianze, et al.
Veröffentlicht: (2026)
von: Zhu, Tianze, et al.
Veröffentlicht: (2026)
Off-Policy Primal-Dual Safe Reinforcement Learning
von: Wu, Zifan, et al.
Veröffentlicht: (2024)
von: Wu, Zifan, et al.
Veröffentlicht: (2024)
A Spectral Revisit of the Distributional Bellman Operator under the Cramér Metric
von: Wang, Keru, et al.
Veröffentlicht: (2026)
von: Wang, Keru, et al.
Veröffentlicht: (2026)
Near-Optimal Sample Complexities of Divergence-based S-rectangular Distributionally Robust Reinforcement Learning
von: Li, Zhenghao, et al.
Veröffentlicht: (2025)
von: Li, Zhenghao, et al.
Veröffentlicht: (2025)
Implicit Safe Set Algorithm for Provably Safe Reinforcement Learning
von: Zhao, Weiye, et al.
Veröffentlicht: (2024)
von: Zhao, Weiye, et al.
Veröffentlicht: (2024)
Policy Constraint by Only Support Constraint for Offline Reinforcement Learning
von: Gao, Yunkai, et al.
Veröffentlicht: (2025)
von: Gao, Yunkai, et al.
Veröffentlicht: (2025)
Constrained Policy Optimization with Explicit Behavior Density for Offline Reinforcement Learning
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
von: Zhang, Jing, et al.
Veröffentlicht: (2023)
Distributional Soft Actor-Critic with Diffusion Policy
von: Liu, Tong, et al.
Veröffentlicht: (2025)
von: Liu, Tong, et al.
Veröffentlicht: (2025)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark
von: Ji, Jiaming, et al.
Veröffentlicht: (2023)
von: Ji, Jiaming, et al.
Veröffentlicht: (2023)
Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
von: Ji, Jiaming, et al.
Veröffentlicht: (2025)
PIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement Learning
von: Huang, Dongchi, et al.
Veröffentlicht: (2025)
von: Huang, Dongchi, et al.
Veröffentlicht: (2025)
Off-Policy Safe Reinforcement Learning with Constrained Optimistic Exploration
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
von: Li, Guopeng, et al.
Veröffentlicht: (2026)
Bootstrap Off-policy with World Model
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
von: Anisimov, Maksim, et al.
Veröffentlicht: (2026)
Revisiting Generative Policies: A Simpler Reinforcement Learning Algorithmic Perspective
von: Zhang, Jinouwen, et al.
Veröffentlicht: (2024)
von: Zhang, Jinouwen, et al.
Veröffentlicht: (2024)
One Filters All: A Generalist Filter for State Estimation
von: Liu, Shiqi, et al.
Veröffentlicht: (2025)
von: Liu, Shiqi, et al.
Veröffentlicht: (2025)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
Safe In-Context Reinforcement Learning
von: Moeini, Amir, et al.
Veröffentlicht: (2025)
von: Moeini, Amir, et al.
Veröffentlicht: (2025)
Enhance the Safety in Reinforcement Learning by ADRC Lagrangian Methods
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
von: Zhang, Mingxu, et al.
Veröffentlicht: (2026)
SafeOR-Gym: A Benchmark Suite for Safe Reinforcement Learning Algorithms on Practical Operations Research Problems
von: Ramanujam, Asha, et al.
Veröffentlicht: (2025)
von: Ramanujam, Asha, et al.
Veröffentlicht: (2025)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
Counterfactually Safe Reinforcement Learning
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
von: Li, Jingyi, et al.
Veröffentlicht: (2026)
Policy Improvement Reinforcement Learning
von: Wang, Huaiyang, et al.
Veröffentlicht: (2026)
von: Wang, Huaiyang, et al.
Veröffentlicht: (2026)
Diffusion-based Reinforcement Learning via Q-weighted Variational Policy Optimization
von: Ding, Shutong, et al.
Veröffentlicht: (2024)
von: Ding, Shutong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning
von: Zhang, Jiaming, et al.
Veröffentlicht: (2026) -
Feasible Policy Iteration for Safe Reinforcement Learning
von: Yang, Yujie, et al.
Veröffentlicht: (2023) -
Policy Bifurcation in Safe Reinforcement Learning
von: Zou, Wenjun, et al.
Veröffentlicht: (2024) -
On the Equilibrium between Feasible Zone and Uncertain Model in Safe Exploration
von: Yang, Yujie, et al.
Veröffentlicht: (2026) -
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
von: Zheng, Yinan, et al.
Veröffentlicht: (2024)