Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Kihyuk, Chae, Woojin, Zhang, Yufan, Lee, Dabeen, Tewari, Ambuj |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025)
by: Hong, Kihyuk, et al.
Published: (2025)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
The Plug-in Approach for Average-Reward and Discounted MDPs: Optimal Sample Complexity Analysis
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Offline Constrained Reinforcement Learning under Partial Data Coverage
by: Ko, Seokmin, et al.
Published: (2025)
by: Ko, Seokmin, et al.
Published: (2025)
Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation
by: Panda, Prashansa, et al.
Published: (2024)
by: Panda, Prashansa, et al.
Published: (2024)
Optimal Horizon-Free Reward-Free Exploration for Linear Mixture MDPs
by: Zhang, Junkai, et al.
Published: (2023)
by: Zhang, Junkai, et al.
Published: (2023)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Infinite-Horizon Reinforcement Learning with Multinomial Logistic Function Approximation
by: Park, Jaehyun, et al.
Published: (2024)
by: Park, Jaehyun, et al.
Published: (2024)
Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount Factor
by: Grand-Clément, Julien, et al.
Published: (2023)
by: Grand-Clément, Julien, et al.
Published: (2023)
When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPs
by: Escamilla, Jose Efraim Aguilar, et al.
Published: (2026)
by: Escamilla, Jose Efraim Aguilar, et al.
Published: (2026)
Achieving Tractable Minimax Optimal Regret in Average Reward MDPs
by: Boone, Victor, et al.
Published: (2024)
by: Boone, Victor, et al.
Published: (2024)
Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach
by: Ganesh, Swetha, et al.
Published: (2025)
by: Ganesh, Swetha, et al.
Published: (2025)
Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
by: Wei, Yukuan, et al.
Published: (2025)
by: Wei, Yukuan, et al.
Published: (2025)
Regret Analysis of Unichain Average Reward Constrained MDPs with General Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Imitation Learning in Discounted Linear MDPs without exploration assumptions
by: Viano, Luca, et al.
Published: (2024)
by: Viano, Luca, et al.
Published: (2024)
Span-Based Optimal Sample Complexity for Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2023)
by: Zurek, Matthew, et al.
Published: (2023)
Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs
by: Thoppe, Gugan, et al.
Published: (2026)
by: Thoppe, Gugan, et al.
Published: (2026)
Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation
by: He, Jianliang, et al.
Published: (2024)
by: He, Jianliang, et al.
Published: (2024)
Learning Weakly Communicating Average-Reward CMDPs: Strong Duality and Improved Regret
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
by: Wang, Shengbo, et al.
Published: (2026)
by: Wang, Shengbo, et al.
Published: (2026)
Global Convergence of Average Reward Constrained MDPs with Neural Critic and General Policy Parameterization
by: Satheesh, Anirudh, et al.
Published: (2026)
by: Satheesh, Anirudh, et al.
Published: (2026)
Online Infinite-Dimensional Regression: Learning Linear Operators
by: Raman, Vinod, et al.
Published: (2023)
by: Raman, Vinod, et al.
Published: (2023)
Policy Zooming: Adaptive Discretization-based Infinite-Horizon Average-Reward Reinforcement Learning
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
by: Lee, Jongmin, et al.
Published: (2025)
by: Lee, Jongmin, et al.
Published: (2025)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs
by: Omidi, Saber, et al.
Published: (2025)
by: Omidi, Saber, et al.
Published: (2025)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
by: Gadot, Uri, et al.
Published: (2023)
by: Gadot, Uri, et al.
Published: (2023)
Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs
by: Zurek, Matthew, et al.
Published: (2024)
by: Zurek, Matthew, et al.
Published: (2024)
Near-Optimal Primal-Dual Algorithm for Learning Linear Mixture CMDPs with Adversarial Rewards
by: Yu, Kihyun, et al.
Published: (2026)
by: Yu, Kihyun, et al.
Published: (2026)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Improved Regret Bound for Safe Reinforcement Learning via Tighter Cost Pessimism and Reward Optimism
by: Yu, Kihyun, et al.
Published: (2024)
by: Yu, Kihyun, et al.
Published: (2024)
Offline-Online Reinforcement Learning for Linear Mixture MDPs
by: Zhang, Zhongjun, et al.
Published: (2026)
by: Zhang, Zhongjun, et al.
Published: (2026)
Efficiently Solving Discounted MDPs with Predictions on Transition Matrices
by: Lyu, Lixing, et al.
Published: (2025)
by: Lyu, Lixing, et al.
Published: (2025)
Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
by: Brooks, Marc, et al.
Published: (2025)
by: Brooks, Marc, et al.
Published: (2025)
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
by: Zamir, Guy, et al.
Published: (2026)
by: Zamir, Guy, et al.
Published: (2026)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
by: Shrestha, Aayam, et al.
Published: (2020)
by: Shrestha, Aayam, et al.
Published: (2020)
Addressing Finite-Horizon MDPs via Low-Rank Tensor Value Approximation
by: Rozada, Sergio, et al.
Published: (2025)
by: Rozada, Sergio, et al.
Published: (2025)
Similar Items
-
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024) -
A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2025) -
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024) -
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024) -
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)