On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wan, Yi, Yu, Huizhen, Sutton, Richard S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
von: Yu, Huizhen, et al.
Veröffentlicht: (2024)
von: Yu, Huizhen, et al.
Veröffentlicht: (2024)
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
von: Yu, Huizhen, et al.
Veröffentlicht: (2025)
von: Yu, Huizhen, et al.
Veröffentlicht: (2025)
Optimal Sample Complexity for Average Reward Markov Decision Processes
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays
von: Yu, Huizhen, et al.
Veröffentlicht: (2023)
von: Yu, Huizhen, et al.
Veröffentlicht: (2023)
Weakly Time-Coupled Approximation of Markov Decision Processes
von: Soheili, Negar, et al.
Veröffentlicht: (2026)
von: Soheili, Negar, et al.
Veröffentlicht: (2026)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
Non-stationary and Varying-discounting Markov Decision Processes for Reinforcement Learning
von: Chen, Zhizuo, et al.
Veröffentlicht: (2025)
von: Chen, Zhizuo, et al.
Veröffentlicht: (2025)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
von: Kumar, Navdeep, et al.
Veröffentlicht: (2024)
Online Reinforcement Learning in Markov Decision Process Using Linear Programming
von: Leon, Vincent, et al.
Veröffentlicht: (2023)
von: Leon, Vincent, et al.
Veröffentlicht: (2023)
Safe Reinforcement Learning for Constrained Markov Decision Processes with Stochastic Stopping Time
von: Mazumdar, Abhijit, et al.
Veröffentlicht: (2024)
von: Mazumdar, Abhijit, et al.
Veröffentlicht: (2024)
Risk-sensitive Markov Decision Process and Learning under General Utility Functions
von: Wu, Zhengqi, et al.
Veröffentlicht: (2023)
von: Wu, Zhengqi, et al.
Veröffentlicht: (2023)
Online Markov Decision Processes with Terminal Law Constraints
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2026)
von: Moreno, Bianca Marin, et al.
Veröffentlicht: (2026)
Central Limit Theorems for Asynchronous Averaged Q-Learning
von: Liu, Xingtu
Veröffentlicht: (2025)
von: Liu, Xingtu
Veröffentlicht: (2025)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
Flipping-based Policy for Chance-Constrained Markov Decision Processes
von: Shen, Xun, et al.
Veröffentlicht: (2024)
von: Shen, Xun, et al.
Veröffentlicht: (2024)
Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator
von: Guo, Zhishuai, et al.
Veröffentlicht: (2021)
von: Guo, Zhishuai, et al.
Veröffentlicht: (2021)
Robust Q-Learning under Corrupted Rewards
von: Maity, Sreejeet, et al.
Veröffentlicht: (2024)
von: Maity, Sreejeet, et al.
Veröffentlicht: (2024)
Robust $Q$-learning Algorithm for Markov Decision Processes under Wasserstein Uncertainty
von: Neufeld, Ariel, et al.
Veröffentlicht: (2022)
von: Neufeld, Ariel, et al.
Veröffentlicht: (2022)
Learning Sequential Decisions from Multiple Sources via Group-Robust Markov Decision Processes
von: Xu, Mingyuan, et al.
Veröffentlicht: (2026)
von: Xu, Mingyuan, et al.
Veröffentlicht: (2026)
Achieving Instance-dependent Sample Complexity for Constrained Markov Decision Process
von: Jiang, Jiashuo, et al.
Veröffentlicht: (2024)
von: Jiang, Jiashuo, et al.
Veröffentlicht: (2024)
Efficient Algorithms for Robust Markov Decision Processes with $s$-Rectangular Ambiguity Sets
von: Ho, Chin Pang, et al.
Veröffentlicht: (2026)
von: Ho, Chin Pang, et al.
Veröffentlicht: (2026)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
Convergence and stability of Q-learning in Hierarchical Reinforcement Learning
von: Manenti, Massimiliano, et al.
Veröffentlicht: (2025)
von: Manenti, Massimiliano, et al.
Veröffentlicht: (2025)
Regularized Q-learning through Robust Averaging
von: Schmitt-Förster, Peter, et al.
Veröffentlicht: (2024)
von: Schmitt-Förster, Peter, et al.
Veröffentlicht: (2024)
Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
von: Wang, Shengbo
Veröffentlicht: (2026)
von: Wang, Shengbo
Veröffentlicht: (2026)
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
von: Kitamura, Toshinori, et al.
Veröffentlicht: (2024)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
Bayesian Ambiguity Contraction-based Adaptive Robust Markov Decision Processes for Adversarial Surveillance Missions
von: Choi, Jimin, et al.
Veröffentlicht: (2025)
von: Choi, Jimin, et al.
Veröffentlicht: (2025)
Optimal Non-Asymptotic Rates of Value Iteration for Average-Reward Markov Decision Processes
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Function Approximation for Non-Markov Processes
von: Kara, Ali Devran
Veröffentlicht: (2026)
von: Kara, Ali Devran
Veröffentlicht: (2026)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
von: Zhang, Weitong, et al.
Veröffentlicht: (2021)
von: Zhang, Weitong, et al.
Veröffentlicht: (2021)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
Finite-Time Complexity of Online Primal-Dual Natural Actor-Critic Algorithm for Constrained Markov Decision Processes
von: Zeng, Sihan, et al.
Veröffentlicht: (2021)
von: Zeng, Sihan, et al.
Veröffentlicht: (2021)
Rates of Convergence in the Central Limit Theorem for Markov Chains, with an Application to TD Learning
von: Srikant, R.
Veröffentlicht: (2024)
von: Srikant, R.
Veröffentlicht: (2024)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2024)
Communication Efficient Federated Learning with Linear Convergence on Heterogeneous Data
von: Liu, Jie, et al.
Veröffentlicht: (2025)
von: Liu, Jie, et al.
Veröffentlicht: (2025)
The Sample-Communication Complexity Trade-off in Federated Q-Learning
von: Salgia, Sudeep, et al.
Veröffentlicht: (2024)
von: Salgia, Sudeep, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
von: Yu, Huizhen, et al.
Veröffentlicht: (2024) -
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
von: Yu, Huizhen, et al.
Veröffentlicht: (2025) -
Optimal Sample Complexity for Average Reward Markov Decision Processes
von: Wang, Shengbo, et al.
Veröffentlicht: (2023) -
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
von: Wang, Shengbo, et al.
Veröffentlicht: (2025) -
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)