Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Jongmin, Ryu, Ernest K. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Policy Gradient Algorithms in Average-Reward Multichain MDPs
by: Lee, Jongmin, et al.
Published: (2026)
by: Lee, Jongmin, et al.
Published: (2026)
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Deflated Dynamics Value Iteration
by: Lee, Jongmin, et al.
Published: (2024)
by: Lee, Jongmin, et al.
Published: (2024)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision Processes
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
A Policy Gradient Primal-Dual Algorithm for Constrained MDPs with Uniform PAC Guarantees
by: Kitamura, Toshinori, et al.
Published: (2024)
by: Kitamura, Toshinori, et al.
Published: (2024)
Policy Gradient in Robust MDPs with Global Convergence Guarantee
by: Wang, Qiuhao, et al.
Published: (2022)
by: Wang, Qiuhao, et al.
Published: (2022)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
by: Zhang, Runyu, et al.
Published: (2023)
by: Zhang, Runyu, et al.
Published: (2023)
Towards Principled, Practical Policy Gradient for Bandits and Tabular MDPs
by: Lu, Michael, et al.
Published: (2024)
by: Lu, Michael, et al.
Published: (2024)
Gradient Starvation in Binary-Reward GRPO: Why Group-Mean Centering Fails and Why the Simplest Fix Works
by: Nie, Wenhua, et al.
Published: (2026)
by: Nie, Wenhua, et al.
Published: (2026)
Policy Gradient Algorithms for Robust MDPs with Non-Rectangular Uncertainty Sets
by: Li, Mengmeng, et al.
Published: (2023)
by: Li, Mengmeng, et al.
Published: (2023)
Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Convergence of Natural Policy Gradient for a Family of Infinite-State Queueing MDPs
by: Grosof, Isaac, et al.
Published: (2024)
by: Grosof, Isaac, et al.
Published: (2024)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Confident Natural Policy Gradient for Local Planning in $q_π$-realizable Constrained MDPs
by: Tian, Tian, et al.
Published: (2024)
by: Tian, Tian, et al.
Published: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Gradient-free Decoder Inversion in Latent Diffusion Models
by: Hong, Seongmin, et al.
Published: (2024)
by: Hong, Seongmin, et al.
Published: (2024)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs
by: Escamilla, Jose Efraim Aguilar, et al.
Published: (2026)
by: Escamilla, Jose Efraim Aguilar, et al.
Published: (2026)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
ETGL-DDPG: A Deep Deterministic Policy Gradient Algorithm for Sparse Reward Continuous Control
by: Futuhi, Ehsan, et al.
Published: (2024)
by: Futuhi, Ehsan, et al.
Published: (2024)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2023)
by: Ding, Dongsheng, et al.
Published: (2023)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025)
by: Wei, Yukuan, et al.
Published: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation
by: Panda, Prashansa, et al.
Published: (2024)
by: Panda, Prashansa, et al.
Published: (2024)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
by: Gadot, Uri, et al.
Published: (2023)
by: Gadot, Uri, et al.
Published: (2023)
LoRA Training in the NTK Regime has No Spurious Local Minima
by: Jang, Uijeong, et al.
Published: (2024)
by: Jang, Uijeong, et al.
Published: (2024)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
by: Bai, Qinbo, et al.
Published: (2023)
by: Bai, Qinbo, et al.
Published: (2023)
LoRA Training Provably Converges to a Low-Rank Global Minimum or It Fails Loudly (But it Probably Won't Fail)
by: Kim, Junsu, et al.
Published: (2025)
by: Kim, Junsu, et al.
Published: (2025)
Natural Policy Gradient for Average Reward Non-Stationary RL
by: Jali, Neharika, et al.
Published: (2025)
by: Jali, Neharika, et al.
Published: (2025)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2025)
by: Ganesh, Swetha, et al.
Published: (2025)
Residual Policy Gradient: A Reward View of KL-regularized Objective
by: Wang, Pengcheng, et al.
Published: (2025)
by: Wang, Pengcheng, et al.
Published: (2025)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs
by: Zhong, Han, et al.
Published: (2021)
by: Zhong, Han, et al.
Published: (2021)
Difference Rewards Policy Gradients
by: Castellini, Jacopo, et al.
Published: (2020)
by: Castellini, Jacopo, et al.
Published: (2020)
Slowly Changing Adversarial Bandit Algorithms are Efficient for Discounted MDPs
by: Kash, Ian A., et al.
Published: (2022)
by: Kash, Ian A., et al.
Published: (2022)
Similar Items
-
Policy Gradient Algorithms in Average-Reward Multichain MDPs
by: Lee, Jongmin, et al.
Published: (2026) -
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
by: Lee, Jongmin, et al.
Published: (2025) -
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024) -
Deflated Dynamics Value Iteration
by: Lee, Jongmin, et al.
Published: (2024) -
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)