On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Navdeep, Murthy, Yashaswini, Shufaro, Itai, Levy, Kfir Y., Srikant, R., Mannor, Shie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2023)
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2023)
The Value of Mechanistic Priors in Sequential Decision Making
von: Shufaro, Itai, et al.
Veröffentlicht: (2026)
von: Shufaro, Itai, et al.
Veröffentlicht: (2026)
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025)
On the Gaussian Limit of the Output of IIR Filters
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2025)
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2025)
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Rates of Convergence in the Central Limit Theorem for Markov Chains, with an Application to TD Learning
von: Srikant, R.
Veröffentlicht: (2024)
von: Srikant, R.
Veröffentlicht: (2024)
Performance of NPG in Countable State-Space Average-Cost RL
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2024)
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2024)
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
von: Koren, Uri, et al.
Veröffentlicht: (2025)
von: Koren, Uri, et al.
Veröffentlicht: (2025)
On Bits and Bandits: Quantifying the Regret-Information Trade-off
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
von: Shufaro, Itai, et al.
Veröffentlicht: (2024)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
von: Wang, Kaixin, et al.
Veröffentlicht: (2023)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
von: Gadot, Uri, et al.
Veröffentlicht: (2023)
Concentration of Cumulative Reward in Markov Decision Processes
von: Sayedana, Borna, et al.
Veröffentlicht: (2024)
von: Sayedana, Borna, et al.
Veröffentlicht: (2024)
An Online Multiobjective Policy Gradient for Long-run Average-reward Markov Decision Process
von: Misra, Rahul, et al.
Veröffentlicht: (2025)
von: Misra, Rahul, et al.
Veröffentlicht: (2025)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
von: Wan, Yi, et al.
Veröffentlicht: (2024)
von: Wan, Yi, et al.
Veröffentlicht: (2024)
Convergence of Distributionally Robust Q-Learning with Linear Function Approximation
von: Mandal, Saptarshi, et al.
Veröffentlicht: (2025)
von: Mandal, Saptarshi, et al.
Veröffentlicht: (2025)
Policy Gradient Methods for Information-Theoretic Opacity in Markov Decision Processes
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
von: Shi, Chongyang, et al.
Veröffentlicht: (2025)
Representative Action Selection for Large Action Space: From Bandits to MDPs
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
Differentially Private Reward Functions in Policy Synthesis for Markov Decision Processes
von: Benvenuti, Alexander, et al.
Veröffentlicht: (2023)
von: Benvenuti, Alexander, et al.
Veröffentlicht: (2023)
Conformal Off-Policy Evaluation in Markov Decision Processes
von: Foffano, Daniele, et al.
Veröffentlicht: (2023)
von: Foffano, Daniele, et al.
Veröffentlicht: (2023)
Optimal Sample Complexity for Average Reward Markov Decision Processes
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
Dual Pricing to Prioritize Renewable Energy and Consumer Preferences in Electricity Markets
von: Jong, Emilie, et al.
Veröffentlicht: (2024)
von: Jong, Emilie, et al.
Veröffentlicht: (2024)
Linear Convergence of Entropy-Regularized Natural Policy Gradient with Linear Function Approximation
von: Cayci, Semih, et al.
Veröffentlicht: (2021)
von: Cayci, Semih, et al.
Veröffentlicht: (2021)
Bayesian Learning of Optimal Policies in Markov Decision Processes with Countably Infinite State-Space
von: Adler, Saghar, et al.
Veröffentlicht: (2023)
von: Adler, Saghar, et al.
Veröffentlicht: (2023)
Linear Convergence of Independent Natural Policy Gradient in Games with Entropy Regularization
von: Sun, Youbang, et al.
Veröffentlicht: (2024)
von: Sun, Youbang, et al.
Veröffentlicht: (2024)
Learning Multiple Initial Solutions to Optimization Problems
von: Sharony, Elad, et al.
Veröffentlicht: (2024)
von: Sharony, Elad, et al.
Veröffentlicht: (2024)
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2026)
An Offline Risk-aware Policy Selection Method for Bayesian Markov Decision Processes
von: Angelotti, Giorgio, et al.
Veröffentlicht: (2021)
von: Angelotti, Giorgio, et al.
Veröffentlicht: (2021)
Representative Action Selection for Large Action Space Bandit Families
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
von: Zhou, Quan, et al.
Veröffentlicht: (2025)
Online Reinforcement Learning in Markov Decision Process Using Linear Programming
von: Leon, Vincent, et al.
Veröffentlicht: (2023)
von: Leon, Vincent, et al.
Veröffentlicht: (2023)
Computing the Exact Pareto Front in Average-Cost Multi-Objective Markov Decision Processes
von: Luo, Jiping, et al.
Veröffentlicht: (2026)
von: Luo, Jiping, et al.
Veröffentlicht: (2026)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
von: Cohen, Lior, et al.
Veröffentlicht: (2026)
Last-Iterate Convergent Policy Gradient Primal-Dual Methods for Constrained MDPs
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
von: Ding, Dongsheng, et al.
Veröffentlicht: (2023)
Approximate Linear Programming for Decentralized Policy Iteration in Cooperative Multi-agent Markov Decision Processes
von: Mandal, Lakshmi, et al.
Veröffentlicht: (2023)
von: Mandal, Lakshmi, et al.
Veröffentlicht: (2023)
OCMDP: Observation-Constrained Markov Decision Process
von: Wang, Taiyi, et al.
Veröffentlicht: (2024)
von: Wang, Taiyi, et al.
Veröffentlicht: (2024)
Stabilizing Policy Gradient Methods via Reward Profiling
von: Ahmed, Shihab, et al.
Veröffentlicht: (2025)
von: Ahmed, Shihab, et al.
Veröffentlicht: (2025)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
Bayesian Ambiguity Contraction-based Adaptive Robust Markov Decision Processes for Adversarial Surveillance Missions
von: Choi, Jimin, et al.
Veröffentlicht: (2025)
von: Choi, Jimin, et al.
Veröffentlicht: (2025)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
von: Du, Yihan, et al.
Veröffentlicht: (2024)
von: Du, Yihan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Convergence of Modified Policy Iteration in Risk Sensitive Exponential Cost Markov Decision Processes
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2023) -
The Value of Mechanistic Priors in Sequential Decision Making
von: Shufaro, Itai, et al.
Veröffentlicht: (2026) -
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2025) -
On the Gaussian Limit of the Output of IIR Filters
von: Murthy, Yashaswini, et al.
Veröffentlicht: (2025) -
On the Convergence of Single-Timescale Actor-Critic
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)