Stabilizing Policy Gradient Methods via Reward Profiling
Fuente:
arXiv
Saved in:
| Main Authors: | Ahmed, Shihab, Bergou, El Houcine, Dutta, Aritra, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
by: Lyu, Xubo, et al.
Published: (2020)
by: Lyu, Xubo, et al.
Published: (2020)
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
by: Vasan, Gautham, et al.
Published: (2024)
by: Vasan, Gautham, et al.
Published: (2024)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022)
by: Bai, Qinbo, et al.
Published: (2022)
Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient
by: Zhang, Xiangyuan, et al.
Published: (2023)
by: Zhang, Xiangyuan, et al.
Published: (2023)
Joint Optimization of Multi-Objective Reinforcement Learning with Policy Gradient Based Algorithm
by: Bai, Qinbo, et al.
Published: (2021)
by: Bai, Qinbo, et al.
Published: (2021)
Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Achieving Tighter Finite-Time Rates for Heterogeneous Federated Stochastic Approximation under Markovian Sampling
by: Zhu, Feng, et al.
Published: (2025)
by: Zhu, Feng, et al.
Published: (2025)
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning
by: Mitra, Aritra, et al.
Published: (2023)
by: Mitra, Aritra, et al.
Published: (2023)
Stability of Primal-Dual Gradient Flow Dynamics for Multi-Block Convex Optimization Problems
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
Model-Free Output Feedback Stabilization via Policy Gradient Methods
by: Zhang, Ankang, et al.
Published: (2026)
by: Zhang, Ankang, et al.
Published: (2026)
Epidemic Control on a Large-Scale-Agent-Based Epidemiology Model using Deep Deterministic Policy Gradient
by: Deshkar, Gaurav, et al.
Published: (2023)
by: Deshkar, Gaurav, et al.
Published: (2023)
Hierarchical Policy-Gradient Reinforcement Learning for Multi-Agent Shepherding Control of Non-Cohesive Targets
by: Covone, Stefano, et al.
Published: (2025)
by: Covone, Stefano, et al.
Published: (2025)
An Offline Risk-aware Policy Selection Method for Bayesian Markov Decision Processes
by: Angelotti, Giorgio, et al.
Published: (2021)
by: Angelotti, Giorgio, et al.
Published: (2021)
STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
by: Gu, Chengyang, et al.
Published: (2026)
by: Gu, Chengyang, et al.
Published: (2026)
Together We Rise: Optimizing Real-Time Multi-Robot Task Allocation using Coordinated Heterogeneous Plays
by: Pal, Aritra, et al.
Published: (2025)
by: Pal, Aritra, et al.
Published: (2025)
Just Few States are Enough: Randomized Sparse Feedback for Stability of Dynamical Systems
by: Hadach, Zaid, et al.
Published: (2025)
by: Hadach, Zaid, et al.
Published: (2025)
Mutual Information as Intrinsic Reward of Reinforcement Learning Agents for On-demand Ride Pooling
by: Zhang, Xianjie, et al.
Published: (2023)
by: Zhang, Xianjie, et al.
Published: (2023)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
CORL: Reinforcement Learning of MILP Policies Solved via Branch and Bound
by: Anand, Akhil S, et al.
Published: (2025)
by: Anand, Akhil S, et al.
Published: (2025)
MathBode: Measuring the Stability of LLM Reasoning using Frequency Response
by: Wang, Charles L.
Published: (2025)
by: Wang, Charles L.
Published: (2025)
Scalable and Interpretable Verification of Image-based Neural Network Controllers for Autonomous Vehicles
by: Parameshwaran, Aditya, et al.
Published: (2025)
by: Parameshwaran, Aditya, et al.
Published: (2025)
CLIP-RLDrive: Human-Aligned Autonomous Driving via CLIP-Based Reward Shaping in Reinforcement Learning
by: Doroudian, Erfan, et al.
Published: (2024)
by: Doroudian, Erfan, et al.
Published: (2024)
Benchmarking Reinforcement Learning via Stochastic Converse Optimality: Generating Systems with Known Optimal Policies
by: Ibrahim, Sinan, et al.
Published: (2026)
by: Ibrahim, Sinan, et al.
Published: (2026)
Efficient On-policy Visual-RL via Stochastic Decoupled Policy Gradient
by: You, Haoxiang, et al.
Published: (2026)
by: You, Haoxiang, et al.
Published: (2026)
Large Language Model Guided Incentive Aware Reward Design for Cooperative Multi-Agent Reinforcement Learning
by: Urgun, Dogan, et al.
Published: (2026)
by: Urgun, Dogan, et al.
Published: (2026)
DiffLight: A Partial Rewards Conditioned Diffusion Model for Traffic Signal Control with Missing Data
by: Chen, Hanyang, et al.
Published: (2024)
by: Chen, Hanyang, et al.
Published: (2024)
Application of Zone Method based Physics-Informed Neural Networks in Reheating Furnaces
by: Dutta, Ujjal Kr, et al.
Published: (2023)
by: Dutta, Ujjal Kr, et al.
Published: (2023)
Operational Wind Speed Forecasts for Chile's Electric Power Sector Using a Hybrid ML Model
by: Suri, Dhruv, et al.
Published: (2024)
by: Suri, Dhruv, et al.
Published: (2024)
DCoPilot: Generative AI-Empowered Policy Adaptation for Dynamic Data Center Operations
by: Li, Minghao, et al.
Published: (2026)
by: Li, Minghao, et al.
Published: (2026)
Robust Q-Learning under Corrupted Rewards
by: Maity, Sreejeet, et al.
Published: (2024)
by: Maity, Sreejeet, et al.
Published: (2024)
An Optimal Policy for Learning Controllable Dynamics by Exploration
by: Loxley, Peter N.
Published: (2025)
by: Loxley, Peter N.
Published: (2025)
Policy Optimization Algorithms in a Unified Framework
by: Wu, Shuang
Published: (2025)
by: Wu, Shuang
Published: (2025)
Certifiably Robust Policies for Uncertain Parametric Environments
by: Schnitzer, Yannik, et al.
Published: (2024)
by: Schnitzer, Yannik, et al.
Published: (2024)
Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization
by: Honari, Homayoun, et al.
Published: (2024)
by: Honari, Homayoun, et al.
Published: (2024)
The Economic Dispatch of Power-to-Gas Systems with Deep Reinforcement Learning:Tackling the Challenge of Delayed Rewards with Long-Term Energy Storage
by: Sage, Manuel, et al.
Published: (2025)
by: Sage, Manuel, et al.
Published: (2025)
Stochastic Approximation with Delayed Updates: Finite-Time Rates under Markovian Sampling
by: Adibi, Arman, et al.
Published: (2024)
by: Adibi, Arman, et al.
Published: (2024)
Conformal Off-Policy Evaluation in Markov Decision Processes
by: Foffano, Daniele, et al.
Published: (2023)
by: Foffano, Daniele, et al.
Published: (2023)
Stabilizing reinforcement learning control: A modular framework for optimizing over all stable behavior
by: Lawrence, Nathan P., et al.
Published: (2023)
by: Lawrence, Nathan P., et al.
Published: (2023)
Exploiting Symmetry in Dynamics for Model-Based Reinforcement Learning with Asymmetric Rewards
by: Sonmez, Yasin, et al.
Published: (2024)
by: Sonmez, Yasin, et al.
Published: (2024)
Reinforcement Learning-enabled Satellite Constellation Reconfiguration and Retasking for Mission-Critical Applications
by: Alami, Hassan El, et al.
Published: (2024)
by: Alami, Hassan El, et al.
Published: (2024)
Similar Items
-
Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
by: Lyu, Xubo, et al.
Published: (2020) -
Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers
by: Vasan, Gautham, et al.
Published: (2024) -
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
by: Bai, Qinbo, et al.
Published: (2022) -
Revisiting LQR Control from the Perspective of Receding-Horizon Policy Gradient
by: Zhang, Xiangyuan, et al.
Published: (2023) -
Joint Optimization of Multi-Objective Reinforcement Learning with Policy Gradient Based Algorithm
by: Bai, Qinbo, et al.
Published: (2021)