A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Kihyuk, Tewari, Ambuj |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024)
by: Hong, Kihyuk, et al.
Published: (2024)
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024)
by: Bai, Qinbo, et al.
Published: (2024)
Compute Aligned Training: Optimizing for Test Time Inference
by: Ousherovitch, Adam, et al.
Published: (2026)
by: Ousherovitch, Adam, et al.
Published: (2026)
Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes
by: Bai, Qinbo, et al.
Published: (2023)
by: Bai, Qinbo, et al.
Published: (2023)
Leveraging Offline Data in Linear Latent Contextual Bandits
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
A Theoretical Framework for Partially Observed Reward-States in RLHF
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
Offline Constrained Reinforcement Learning under Partial Data Coverage
by: Ko, Seokmin, et al.
Published: (2025)
by: Ko, Seokmin, et al.
Published: (2025)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Quantum Speedups in Regret Analysis of Infinite Horizon Average-Reward Markov Decision Processes
by: Ganguly, Bhargav, et al.
Published: (2023)
by: Ganguly, Bhargav, et al.
Published: (2023)
Efficient Online Learning with Offline Datasets for Infinite Horizon MDPs: A Bayesian Approach
by: Tang, Dengwang, et al.
Published: (2023)
by: Tang, Dengwang, et al.
Published: (2023)
Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions
by: Brooks, Marc, et al.
Published: (2025)
by: Brooks, Marc, et al.
Published: (2025)
Online Infinite-Dimensional Regression: Learning Linear Operators
by: Raman, Vinod, et al.
Published: (2023)
by: Raman, Vinod, et al.
Published: (2023)
Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs
by: Ganesh, Swetha, et al.
Published: (2024)
by: Ganesh, Swetha, et al.
Published: (2024)
Improved Sample Complexity Analysis of Natural Policy Gradient Algorithm with General Parameterization for Infinite Horizon Discounted Reward Markov Decision Processes
by: Mondal, Washim Uddin, et al.
Published: (2023)
by: Mondal, Washim Uddin, et al.
Published: (2023)
Hierarchical Average-Reward Linearly-solvable Markov Decision Processes
by: Infante, Guillermo, et al.
Published: (2024)
by: Infante, Guillermo, et al.
Published: (2024)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
A Characterization of List Language Identification in the Limit
by: Charikar, Moses, et al.
Published: (2025)
by: Charikar, Moses, et al.
Published: (2025)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
DeepAveragers: Offline Reinforcement Learning by Solving Derived Non-Parametric MDPs
by: Shrestha, Aayam, et al.
Published: (2020)
by: Shrestha, Aayam, et al.
Published: (2020)
A Training-free Method for LLM Text Attribution
by: Radvand, Tara, et al.
Published: (2025)
by: Radvand, Tara, et al.
Published: (2025)
Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm
by: Vakili, Sattar, et al.
Published: (2024)
by: Vakili, Sattar, et al.
Published: (2024)
Constrained Reinforcement Learning with Average Reward Objective: Model-Based and Model-Free Algorithms
by: Aggarwal, Vaneet, et al.
Published: (2024)
by: Aggarwal, Vaneet, et al.
Published: (2024)
Average-Reward Soft Actor-Critic
by: Adamczyk, Jacob, et al.
Published: (2025)
by: Adamczyk, Jacob, et al.
Published: (2025)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
by: Muni, Aneri, et al.
Published: (2026)
by: Muni, Aneri, et al.
Published: (2026)
Efficient $Q$-Learning and Actor-Critic Methods for Robust Average Reward Reinforcement Learning
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
WARP: On the Benefits of Weight Averaged Rewarded Policies
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
Provably Efficient Infinite-Horizon Average-Reward Reinforcement Learning with Linear Function Approximation
by: Chae, Woojin, et al.
Published: (2024)
by: Chae, Woojin, et al.
Published: (2024)
A Harmonic Mean Formulation of Average Reward Reinforcement Learning in SMDPs
by: Shtossel, Erel, et al.
Published: (2026)
by: Shtossel, Erel, et al.
Published: (2026)
An Asymptotically Optimal Algorithm for the Convex Hull Membership Problem
by: Qiao, Gang, et al.
Published: (2023)
by: Qiao, Gang, et al.
Published: (2023)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
by: Kar, Avik, et al.
Published: (2024)
by: Kar, Avik, et al.
Published: (2024)
Optimal Thresholding Linear Bandit
by: Rivera, Eduardo Ochoa, et al.
Published: (2024)
by: Rivera, Eduardo Ochoa, et al.
Published: (2024)
Graph Neural Diffusion Networks for Semi-supervised Learning
by: Ye, Wei, et al.
Published: (2022)
by: Ye, Wei, et al.
Published: (2022)
If generative AI is the answer, what is the question?
by: Tewari, Ambuj
Published: (2025)
by: Tewari, Ambuj
Published: (2025)
Controlling Statistical, Discretization, and Truncation Errors in Learning Fourier Linear Operators
by: Subedi, Unique, et al.
Published: (2024)
by: Subedi, Unique, et al.
Published: (2024)
WARM: On the Benefits of Weight Averaged Reward Models
by: Ramé, Alexandre, et al.
Published: (2024)
by: Ramé, Alexandre, et al.
Published: (2024)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Similar Items
-
Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs
by: Hong, Kihyuk, et al.
Published: (2024) -
Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span
by: Chae, Woojin, et al.
Published: (2024) -
A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs
by: Hong, Kihyuk, et al.
Published: (2024) -
Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm
by: Bai, Qinbo, et al.
Published: (2024) -
Compute Aligned Training: Optimizing for Test Time Inference
by: Ousherovitch, Adam, et al.
Published: (2026)