Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
Fuente:
arXiv
Saved in:
| Main Authors: | Lyu, Xubo, Li, Site, Siriya, Seth, Pu, Ye, Chen, Mo |
|---|---|
| Format: | Preprint |
| Published: |
2020
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Fast and Safety-Guaranteed Trajectory Planning and Tracking for Time-Varying Systems
by: Siriya, Seth, et al.
Published: (2024)
by: Siriya, Seth, et al.
Published: (2024)
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
by: Vasan, Gautham, et al.
Published: (2024)
by: Vasan, Gautham, et al.
Published: (2024)
An Optimal Policy for Learning Controllable Dynamics by Exploration
by: Loxley, Peter N.
Published: (2025)
by: Loxley, Peter N.
Published: (2025)
Embedding Morphology into Transformers for Cross-Robot Policy Learning
by: Suzuki, Kei, et al.
Published: (2026)
by: Suzuki, Kei, et al.
Published: (2026)
Scaling Learning based Policy Optimization for Temporal Logic Tasks by Controller Network Dropout
by: Hashemi, Navid, et al.
Published: (2024)
by: Hashemi, Navid, et al.
Published: (2024)
Non-Asymptotic Bounds for Closed-Loop Identification of Unstable Nonlinear Stochastic Systems
by: Siriya, Seth, et al.
Published: (2024)
by: Siriya, Seth, et al.
Published: (2024)
A Framework for Adaptive Stabilisation of Nonlinear Stochastic Systems
by: Siriya, Seth, et al.
Published: (2025)
by: Siriya, Seth, et al.
Published: (2025)
Hybrid Quantum-Classical Policy Gradient for Adaptive Control of Cyber-Physical Systems: A Comparative Study of VQC vs. MLP
by: Aueawatthanaphisut, Aueaphum, et al.
Published: (2025)
by: Aueawatthanaphisut, Aueaphum, et al.
Published: (2025)
Constrained Control for Autonomous Spacecraft Rendezvous: Learning-Based Time Shift Governor
by: Kim, Taehyeun, et al.
Published: (2024)
by: Kim, Taehyeun, et al.
Published: (2024)
SpaceMind: A Modular and Self-Evolving Embodied Vision-Language Agent Framework for Autonomous On-orbit Servicing
by: Wu, Aodi, et al.
Published: (2026)
by: Wu, Aodi, et al.
Published: (2026)
Align and Filter: Improving Performance in Asynchronous On-Policy RL
by: Honari, Homayoun, et al.
Published: (2026)
by: Honari, Homayoun, et al.
Published: (2026)
Estimating Control Barriers from Offline Data
by: Yu, Hongzhan, et al.
Published: (2025)
by: Yu, Hongzhan, et al.
Published: (2025)
SAPG: Split and Aggregate Policy Gradients
by: Singla, Jayesh, et al.
Published: (2024)
by: Singla, Jayesh, et al.
Published: (2024)
Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving
by: Zhang, Zhihao, et al.
Published: (2025)
by: Zhang, Zhihao, et al.
Published: (2025)
Predictive Red Teaming: Breaking Policies Without Breaking Robots
by: Majumdar, Anirudha, et al.
Published: (2025)
by: Majumdar, Anirudha, et al.
Published: (2025)
Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization
by: Honari, Homayoun, et al.
Published: (2024)
by: Honari, Homayoun, et al.
Published: (2024)
Explainable Representation of Finite-Memory Policies for POMDPs using Decision Trees
by: Azeem, Muqsit, et al.
Published: (2024)
by: Azeem, Muqsit, et al.
Published: (2024)
Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees
by: Ganguly, Sourav, et al.
Published: (2025)
by: Ganguly, Sourav, et al.
Published: (2025)
Deep Reinforcement Learning for Advanced Longitudinal Control and Collision Avoidance in High-Risk Driving Scenarios
by: Chen, Dianwei, et al.
Published: (2024)
by: Chen, Dianwei, et al.
Published: (2024)
Meta SAC-Lag: Towards Deployable Safe Reinforcement Learning via MetaGradient-based Hyperparameter Tuning
by: Honari, Homayoun, et al.
Published: (2024)
by: Honari, Homayoun, et al.
Published: (2024)
Control-ITRA: Controlling the Behavior of a Driving Model
by: Lioutas, Vasileios, et al.
Published: (2025)
by: Lioutas, Vasileios, et al.
Published: (2025)
ManyQuadrupeds: Learning a Single Locomotion Policy for Diverse Quadruped Robots
by: Shafiee, Milad, et al.
Published: (2023)
by: Shafiee, Milad, et al.
Published: (2023)
Learning to Drift in Extreme Turning with Active Exploration and Gaussian Process Based MPC
by: Wu, Guoqiang, et al.
Published: (2024)
by: Wu, Guoqiang, et al.
Published: (2024)
Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
by: Murillo-Gonzalez, Alejandro, et al.
Published: (2026)
Multi-CALF: A Policy Combination Approach with Statistical Guarantees
by: Malaniya, Georgiy, et al.
Published: (2025)
by: Malaniya, Georgiy, et al.
Published: (2025)
RL + Model-based Control: Using On-demand Optimal Control to Learn Versatile Legged Locomotion
by: Kang, Dongho, et al.
Published: (2023)
by: Kang, Dongho, et al.
Published: (2023)
The Role of Touch: Towards Optimal Tactile Sensing Distribution in Anthropomorphic Hands for Dexterous In-Hand Manipulation
by: Almeida, João Damião, et al.
Published: (2025)
by: Almeida, João Damião, et al.
Published: (2025)
Visual CPG-RL: Learning Central Pattern Generators for Visually-Guided Quadruped Locomotion
by: Bellegarda, Guillaume, et al.
Published: (2022)
by: Bellegarda, Guillaume, et al.
Published: (2022)
Learning from Imperfect Demonstrations via Temporal Behavior Tree-Guided Trajectory Repair
by: Puranic, Aniruddh G., et al.
Published: (2026)
by: Puranic, Aniruddh G., et al.
Published: (2026)
Frugal Actor-Critic: Sample Efficient Off-Policy Deep Reinforcement Learning Using Unique Experiences
by: Singh, Nikhil Kumar, et al.
Published: (2024)
by: Singh, Nikhil Kumar, et al.
Published: (2024)
GUIDEd Agents: Enhancing Navigation Policies through Task-Specific Uncertainty Abstraction in Localization-Limited Environments
by: Puthumanaillam, Gokul, et al.
Published: (2024)
by: Puthumanaillam, Gokul, et al.
Published: (2024)
Learning Force Control for Legged Manipulation
by: Portela, Tifanny, et al.
Published: (2024)
by: Portela, Tifanny, et al.
Published: (2024)
Towards a Practical Understanding of Lagrangian Methods in Safe Reinforcement Learning
by: Spoor, Lindsay, et al.
Published: (2025)
by: Spoor, Lindsay, et al.
Published: (2025)
A Safe Reinforcement Learning driven Weights-varying Model Predictive Control for Autonomous Vehicle Motion Control
by: Zarrouki, Baha, et al.
Published: (2024)
by: Zarrouki, Baha, et al.
Published: (2024)
Solving Multi-Agent Safe Optimal Control with Distributed Epigraph Form MARL
by: Zhang, Songyuan, et al.
Published: (2025)
by: Zhang, Songyuan, et al.
Published: (2025)
SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving
by: Wu, Kangyu, et al.
Published: (2026)
by: Wu, Kangyu, et al.
Published: (2026)
Generative Predictive Control: Flow Matching Policies for Dynamic and Difficult-to-Demonstrate Tasks
by: Kurtz, Vince, et al.
Published: (2025)
by: Kurtz, Vince, et al.
Published: (2025)
MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety
by: Wang, Justin, et al.
Published: (2024)
by: Wang, Justin, et al.
Published: (2024)
DexterityGen: Foundation Controller for Unprecedented Dexterity
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
A Physics-Informed Machine Learning Framework for Safe and Optimal Control of Autonomous Systems
by: Tayal, Manan, et al.
Published: (2025)
by: Tayal, Manan, et al.
Published: (2025)
Similar Items
-
Towards Fast and Safety-Guaranteed Trajectory Planning and Tracking for Time-Varying Systems
by: Siriya, Seth, et al.
Published: (2024) -
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
by: Vasan, Gautham, et al.
Published: (2024) -
An Optimal Policy for Learning Controllable Dynamics by Exploration
by: Loxley, Peter N.
Published: (2025) -
Embedding Morphology into Transformers for Cross-Robot Policy Learning
by: Suzuki, Kei, et al.
Published: (2026) -
Scaling Learning based Policy Optimization for Temporal Logic Tasks by Controller Network Dropout
by: Hashemi, Navid, et al.
Published: (2024)