Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chae, Woojin, Hong, Kihyuk, Zhang, Yufan, Tewari, Ambuj, Lee, Dabeen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
von: Chae, Woojin, et al.
Veröffentlicht: (2024)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025)
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
von: Zhang, Junkai, et al.
Veröffentlicht: (2023)
von: Zhang, Junkai, et al.
Veröffentlicht: (2023)
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
von: Park, Jaehyun, et al.
Veröffentlicht: (2024)
von: Park, Jaehyun, et al.
Veröffentlicht: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
von: Yu, Kihyun, et al.
Veröffentlicht: (2026)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024)
Span-Based Optimal Sample Complexity for Average Reward MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
von: Zurek, Matthew, et al.
Veröffentlicht: (2023)
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
von: Yu, Kihyun, et al.
Veröffentlicht: (2024)
von: Yu, Kihyun, et al.
Veröffentlicht: (2024)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
von: Zamir, Guy, et al.
Veröffentlicht: (2026)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
von: Zhang, Zhongjun, et al.
Veröffentlicht: (2026)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
von: Boone, Victor, et al.
Veröffentlicht: (2024)
von: Boone, Victor, et al.
Veröffentlicht: (2024)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
von: Wang, Shengbo, et al.
Veröffentlicht: (2026)
Planning and Learning in Average Risk-aware MDPs
von: Wang, Weikai, et al.
Veröffentlicht: (2025)
von: Wang, Weikai, et al.
Veröffentlicht: (2025)
Stochastic-Constrained Stochastic Optimization with Markovian Data
von: Kim, Yeongjong, et al.
Veröffentlicht: (2023)
von: Kim, Yeongjong, et al.
Veröffentlicht: (2023)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
von: Omidi, Saber, et al.
Veröffentlicht: (2025)
von: Omidi, Saber, et al.
Veröffentlicht: (2025)
Parameter-Free Algorithms for Performative Regret Minimization under Decision-Dependent Distributions
von: Park, Sungwoo, et al.
Veröffentlicht: (2024)
von: Park, Sungwoo, et al.
Veröffentlicht: (2024)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
von: Zurek, Matthew, et al.
Veröffentlicht: (2024)
Span-Agnostic Optimal Sample Complexity and Oracle Inequalities for Average-Reward RL
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
von: Zurek, Matthew, et al.
Veröffentlicht: (2025)
Landscape of Policy Optimization for Finite Horizon MDPs with General State and Action
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
Learning Neural Contracting Dynamics: Extended Linearization and Global Guarantees
von: Jaffe, Sean, et al.
Veröffentlicht: (2024)
von: Jaffe, Sean, et al.
Veröffentlicht: (2024)
Chebyshev Center-Based Direction Selection for Multi-Objective Optimization and Training PINNs
von: Yoon, Hoyeol, et al.
Veröffentlicht: (2026)
von: Yoon, Hoyeol, et al.
Veröffentlicht: (2026)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
von: Chen, Zijun, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Infinite Horizon Mean Field Problems in Continuous Spaces
von: Angiuli, Andrea, et al.
Veröffentlicht: (2023)
von: Angiuli, Andrea, et al.
Veröffentlicht: (2023)
Soft Robust MDPs and Risk-Sensitive MDPs: Equivalence, Policy Gradient, and Sample Complexity
von: Zhang, Runyu, et al.
Veröffentlicht: (2023)
von: Zhang, Runyu, et al.
Veröffentlicht: (2023)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
von: Wan, Yi, et al.
Veröffentlicht: (2024)
von: Wan, Yi, et al.
Veröffentlicht: (2024)
Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
von: Li, Jingqi, et al.
Veröffentlicht: (2022)
von: Li, Jingqi, et al.
Veröffentlicht: (2022)
Constrained Average-Reward Intermittently Observable MDPs
von: Avrachenkov, Konstantin, et al.
Veröffentlicht: (2025)
von: Avrachenkov, Konstantin, et al.
Veröffentlicht: (2025)
Reward-Relevance-Filtered Linear Offline Reinforcement Learning
von: Zhou, Angela
Veröffentlicht: (2024)
von: Zhou, Angela
Veröffentlicht: (2024)
Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
von: Kaya, Ege C., et al.
Veröffentlicht: (2026)
Optimal Sample Complexity for Average Reward Markov Decision Processes
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
von: Wang, Shengbo, et al.
Veröffentlicht: (2023)
Policy Gradient Algorithms in Average-Reward Multichain MDPs
von: Lee, Jongmin, et al.
Veröffentlicht: (2026)
von: Lee, Jongmin, et al.
Veröffentlicht: (2026)
Offline Constrained Reinforcement Learning under Partial Data Coverage
von: Ko, Seokmin, et al.
Veröffentlicht: (2025)
von: Ko, Seokmin, et al.
Veröffentlicht: (2025)
Offline Reinforcement Learning via Linear-Programming with Error-Bound Induced Constraints
von: Ozdaglar, Asuman, et al.
Veröffentlicht: (2022)
von: Ozdaglar, Asuman, et al.
Veröffentlicht: (2022)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
Regret Lower Bounds for Learning Linear Quadratic Gaussian Systems
von: Ziemann, Ingvar, et al.
Veröffentlicht: (2022)
von: Ziemann, Ingvar, et al.
Veröffentlicht: (2022)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
von: Wang, Shengbo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2024) -
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
von: Chae, Woojin, et al.
Veröffentlicht: (2024) -
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
von: Hong, Kihyuk, et al.
Veröffentlicht: (2025) -
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
von: Yu, Kihyun, et al.
Veröffentlicht: (2026) -
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
von: Zhang, Junkai, et al.
Veröffentlicht: (2023)