Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shengbo, Si, Nian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Optimal Sample Complexity for Average Reward Markov Decision Processes
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025)
by: Chen, Zijun, et al.
Published: (2025)
Tractable Robust Markov Decision Processes
by: Grand-Clément, Julien, et al.
Published: (2024)
by: Grand-Clément, Julien, et al.
Published: (2024)
Learning Sequential Decisions from Multiple Sources via Group-Robust Markov Decision Processes
by: Xu, Mingyuan, et al.
Published: (2026)
by: Xu, Mingyuan, et al.
Published: (2026)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
by: Wan, Yi, et al.
Published: (2024)
by: Wan, Yi, et al.
Published: (2024)
On the Foundation of Distributionally Robust Reinforcement Learning
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Sample Complexity of Variance-reduced Distributionally Robust Q-learning
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
by: Kaya, Ege C., et al.
Published: (2026)
by: Kaya, Ege C., et al.
Published: (2026)
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
by: Kitamura, Toshinori, et al.
Published: (2024)
by: Kitamura, Toshinori, et al.
Published: (2024)
Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets
by: Ho, Chin Pang, et al.
Published: (2026)
by: Ho, Chin Pang, et al.
Published: (2026)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision Processes
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Bayesian Ambiguity Contraction-based Adaptive Robust Markov Decision Processes for Adversarial Surveillance Missions
by: Choi, Jimin, et al.
Published: (2025)
by: Choi, Jimin, et al.
Published: (2025)
Weakly Time-Coupled Approximation of Markov Decision Processes
by: Soheili, Negar, et al.
Published: (2026)
by: Soheili, Negar, et al.
Published: (2026)
Online Markov Decision Processes with Terminal Law Constraints
by: Moreno, Bianca Marin, et al.
Published: (2026)
by: Moreno, Bianca Marin, et al.
Published: (2026)
Flipping-based Policy for Chance-Constrained Markov Decision Processes
by: Shen, Xun, et al.
Published: (2024)
by: Shen, Xun, et al.
Published: (2024)
Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
by: Wang, Shengbo
Published: (2026)
by: Wang, Shengbo
Published: (2026)
Span-Based Optimal Sample Complexity for Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2023)
by: Zurek, Matthew, et al.
Published: (2023)
Non-stationary and Varying-discounting Markov Decision Processes for Reinforcement Learning
by: Chen, Zhizuo, et al.
Published: (2025)
by: Chen, Zhizuo, et al.
Published: (2025)
Achieving Instance-dependent Sample Complexity for Constrained Markov Decision Process
by: Jiang, Jiashuo, et al.
Published: (2024)
by: Jiang, Jiashuo, et al.
Published: (2024)
Online Reinforcement Learning in Markov Decision Process Using Linear Programming
by: Leon, Vincent, et al.
Published: (2023)
by: Leon, Vincent, et al.
Published: (2023)
Safe Reinforcement Learning for Constrained Markov Decision Processes with Stochastic Stopping Time
by: Mazumdar, Abhijit, et al.
Published: (2024)
by: Mazumdar, Abhijit, et al.
Published: (2024)
Risk-sensitive Markov Decision Process and Learning under General Utility Functions
by: Wu, Zhengqi, et al.
Published: (2023)
by: Wu, Zhengqi, et al.
Published: (2023)
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
by: Zurek, Matthew, et al.
Published: (2025)
by: Zurek, Matthew, et al.
Published: (2025)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Robust $Q$-learning Algorithm for Markov Decision Processes under Wasserstein Uncertainty
by: Neufeld, Ariel, et al.
Published: (2022)
by: Neufeld, Ariel, et al.
Published: (2022)
Is Bellman Equation Enough for Learning Control?
by: You, Haoxiang, et al.
Published: (2025)
by: You, Haoxiang, et al.
Published: (2025)
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
by: Zurek, Matthew, et al.
Published: (2025)
by: Zurek, Matthew, et al.
Published: (2025)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
by: Zhang, Weitong, et al.
Published: (2021)
by: Zhang, Weitong, et al.
Published: (2021)
Robust Regression over Averaged Uncertainty
by: Bertsimas, Dimitris, et al.
Published: (2023)
by: Bertsimas, Dimitris, et al.
Published: (2023)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Finite-Time Complexity of Online Primal-Dual Natural Actor-Critic Algorithm for Constrained Markov Decision Processes
by: Zeng, Sihan, et al.
Published: (2021)
by: Zeng, Sihan, et al.
Published: (2021)
Regularized Q-learning through Robust Averaging
by: Schmitt-Förster, Peter, et al.
Published: (2024)
by: Schmitt-Förster, Peter, et al.
Published: (2024)
Optimal Parameter Adaptation for Safety-Critical Control via Safe Barrier Bayesian Optimization
by: Wang, Shengbo, et al.
Published: (2025)
by: Wang, Shengbo, et al.
Published: (2025)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025)
by: Omidi, Saber, et al.
Published: (2025)
Similar Items
-
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026) -
Optimal Sample Complexity for Average Reward Markov Decision Processes
by: Wang, Shengbo, et al.
Published: (2023) -
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
by: Chen, Zijun, et al.
Published: (2025) -
Tractable Robust Markov Decision Processes
by: Grand-Clément, Julien, et al.
Published: (2024) -
Learning Sequential Decisions from Multiple Sources via Group-Robust Markov Decision Processes
by: Xu, Mingyuan, et al.
Published: (2026)