Conformal Symplectic Optimization for Stable Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Yao, Zhang, Xiangteng, Li, Shengbo Eben, Duan, Jingliang, Tao, Letian, Xu, Qing, He, Lei, Li, Keqiang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
by: Wang, Likun, et al.
Published: (2025)
by: Wang, Likun, et al.
Published: (2025)
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
by: Liu, Shiqi, et al.
Published: (2026)
by: Liu, Shiqi, et al.
Published: (2026)
Hierarchical End-to-End Autonomous Driving: Integrating BEV Perception with Deep Reinforcement Learning
by: Lu, Siyi, et al.
Published: (2024)
by: Lu, Siyi, et al.
Published: (2024)
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
by: Zhang, Feihong, et al.
Published: (2025)
by: Zhang, Feihong, et al.
Published: (2025)
Bootstrap Off-policy with World Model
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
by: Zhan, Guojian, et al.
Published: (2025)
by: Zhan, Guojian, et al.
Published: (2025)
Policy Bifurcation in Safe Reinforcement Learning
by: Zou, Wenjun, et al.
Published: (2024)
by: Zou, Wenjun, et al.
Published: (2024)
Zeroth-Order Actor-Critic: An Evolutionary Framework for Sequential Decision Problems
by: Lei, Yuheng, et al.
Published: (2022)
by: Lei, Yuheng, et al.
Published: (2022)
Predictive Lagrangian Optimization for Constrained Reinforcement Learning
by: Zhang, Tianqi, et al.
Published: (2025)
by: Zhang, Tianqi, et al.
Published: (2025)
Distributional Soft Actor-Critic with Diffusion Policy
by: Liu, Tong, et al.
Published: (2025)
by: Liu, Tong, et al.
Published: (2025)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
by: Zhan, Guojian, et al.
Published: (2026)
by: Zhan, Guojian, et al.
Published: (2026)
Safe Offline Reinforcement Learning with Feasibility-Guided Diffusion Model
by: Zheng, Yinan, et al.
Published: (2024)
by: Zheng, Yinan, et al.
Published: (2024)
Diffusion Actor-Critic with Entropy Regulator
by: Wang, Yinuo, et al.
Published: (2024)
by: Wang, Yinuo, et al.
Published: (2024)
One Filters All: A Generalist Filter for State Estimation
by: Liu, Shiqi, et al.
Published: (2025)
by: Liu, Shiqi, et al.
Published: (2025)
Enhanced DACER Algorithm with High Diffusion Efficiency
by: Wang, Yinuo, et al.
Published: (2025)
by: Wang, Yinuo, et al.
Published: (2025)
ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization
by: Yang, Letian, et al.
Published: (2026)
by: Yang, Letian, et al.
Published: (2026)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
by: Zhang, Han, et al.
Published: (2025)
by: Zhang, Han, et al.
Published: (2025)
Jump-Start Reinforcement Learning with Self-Evolving Priors for Extreme Monopedal Locomotion
by: Zheng, Ziang, et al.
Published: (2025)
by: Zheng, Ziang, et al.
Published: (2025)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
by: Zhang, Tianle, et al.
Published: (2024)
by: Zhang, Tianle, et al.
Published: (2024)
FedEL: Federated Elastic Learning for Heterogeneous Devices
by: Zhang, Letian, et al.
Published: (2025)
by: Zhang, Letian, et al.
Published: (2025)
Towards Automated Semantic Interpretability in Reinforcement Learning via Vision-Language Models
by: Li, Zhaoxin, et al.
Published: (2025)
by: Li, Zhaoxin, et al.
Published: (2025)
Canonical Form of Datatic Description in Control Systems
by: Zhan, Guojian, et al.
Published: (2024)
by: Zhan, Guojian, et al.
Published: (2024)
FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-Tuning
by: Bian, Jieming, et al.
Published: (2026)
by: Bian, Jieming, et al.
Published: (2026)
Flow-Based Policy for Online Reinforcement Learning
by: Lv, Lei, et al.
Published: (2025)
by: Lv, Lei, et al.
Published: (2025)
Mixed Policy Gradient: off-policy reinforcement learning driven jointly by data and model
by: Guan, Yang, et al.
Published: (2021)
by: Guan, Yang, et al.
Published: (2021)
Distributional Soft Actor-Critic with Three Refinements
by: Duan, Jingliang, et al.
Published: (2023)
by: Duan, Jingliang, et al.
Published: (2023)
Mildly Conservative Q-Learning for Offline Reinforcement Learning
by: Lyu, Jiafei, et al.
Published: (2022)
by: Lyu, Jiafei, et al.
Published: (2022)
EXPO: Stable Reinforcement Learning with Expressive Policies
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
ODRL: A Benchmark for Off-Dynamics Reinforcement Learning
by: Lyu, Jiafei, et al.
Published: (2024)
by: Lyu, Jiafei, et al.
Published: (2024)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
by: Feng, Lang, et al.
Published: (2026)
by: Feng, Lang, et al.
Published: (2026)
Learning to Optimize for Reinforcement Learning
by: Lan, Qingfeng, et al.
Published: (2023)
by: Lan, Qingfeng, et al.
Published: (2023)
Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model
by: Gao, Zeyu, et al.
Published: (2022)
by: Gao, Zeyu, et al.
Published: (2022)
Neural Beam Field for Spatial Beam RSRP Prediction
by: Guo, Keqiang, et al.
Published: (2025)
by: Guo, Keqiang, et al.
Published: (2025)
FedALT: Federated Fine-Tuning through Adaptive Local Training with Rest-of-World LoRA
by: Bian, Jieming, et al.
Published: (2025)
by: Bian, Jieming, et al.
Published: (2025)
Pentest-R1: Towards Autonomous Penetration Testing Reasoning Optimized via Two-Stage Reinforcement Learning
by: Kong, He, et al.
Published: (2025)
by: Kong, He, et al.
Published: (2025)
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
by: Hu, Jifeng, et al.
Published: (2025)
by: Hu, Jifeng, et al.
Published: (2025)
Hierarchical Feature-level Reverse Propagation for Post-Training Neural Networks
by: Ding, Ni, et al.
Published: (2025)
by: Ding, Ni, et al.
Published: (2025)
DeCoR: Design and Control Co-Optimization for Urban Streets Using Reinforcement Learning
by: Poudel, Bibek, et al.
Published: (2026)
by: Poudel, Bibek, et al.
Published: (2026)
Similar Items
-
Off-policy Reinforcement Learning with Model-based Exploration Augmentation
by: Wang, Likun, et al.
Published: (2025) -
Augmented Lagrangian Multiplier Network for State-wise Safety in Reinforcement Learning
by: Zhang, Jiaming, et al.
Published: (2026) -
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
by: Liu, Shiqi, et al.
Published: (2026) -
Hierarchical End-to-End Autonomous Driving: Integrating BEV Perception with Deep Reinforcement Learning
by: Lu, Siyi, et al.
Published: (2024) -
Distributional Soft Actor-Critic with Harmonic Gradient for Safe and Efficient Autonomous Driving in Multi-lane Scenarios
by: Zhang, Feihong, et al.
Published: (2025)