Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Huizhen, Wan, Yi, Sutton, Richard S. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
by: Yu, Huizhen, et al.
Published: (2024)
by: Yu, Huizhen, et al.
Published: (2024)
A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays
by: Yu, Huizhen, et al.
Published: (2023)
by: Yu, Huizhen, et al.
Published: (2023)
Sample Complexity of Policy Gradient for Log-Growth Control
by: Pan, Qiuhua, et al.
Published: (2026)
by: Pan, Qiuhua, et al.
Published: (2026)
Average Cost Optimality of Partially Observed MDPS: Contraction of Non-linear Filters, Optimal Solutions and Approximations
by: Demirci, Yunus Emre, et al.
Published: (2023)
by: Demirci, Yunus Emre, et al.
Published: (2023)
On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies
by: Yu, Huizhen
Published: (2022)
by: Yu, Huizhen
Published: (2022)
Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation
by: Zheng, Yaowei, et al.
Published: (2026)
by: Zheng, Yaowei, et al.
Published: (2026)
Long run control of nonhomogeneous Markov processes
by: Stettner, Łukasz
Published: (2025)
by: Stettner, Łukasz
Published: (2025)
Stochastic Control with Signatures
by: Bank, P., et al.
Published: (2024)
by: Bank, P., et al.
Published: (2024)
A Sequential Testing Problem with Signal Control
by: Campbell, Steven, et al.
Published: (2025)
by: Campbell, Steven, et al.
Published: (2025)
Stability and Sensitivity Analysis of Relative Temporal-Difference Learning: Extended Version
by: Sakha, Masoud S., et al.
Published: (2026)
by: Sakha, Masoud S., et al.
Published: (2026)
Optimistic Training and Convergence of Q-Learning -- Extended Version
by: Mehta, Prashant, et al.
Published: (2026)
by: Mehta, Prashant, et al.
Published: (2026)
Robust Ergodic Control of Jump-Diffusion Systems under Drift and Intensity Uncertainty
by: Azze, Abel, et al.
Published: (2026)
by: Azze, Abel, et al.
Published: (2026)
A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces
by: Kerimkulov, Bekzhan, et al.
Published: (2023)
by: Kerimkulov, Bekzhan, et al.
Published: (2023)
Causal Hamilton-Jacobi-Bellman Equations for Anticipative Stochastic Optimal Control
by: Bank, Peter, et al.
Published: (2025)
by: Bank, Peter, et al.
Published: (2025)
Cost-optimal Management of a Residential Heating System With a Geothermal Energy Storage Under Uncertainty
by: Takam, Paul Honore, et al.
Published: (2025)
by: Takam, Paul Honore, et al.
Published: (2025)
Non-local Hamilton-Jacobi-Bellman equations for the stochastic optimal control of path-dependent piecewise deterministic processes
by: Bandini, Elena, et al.
Published: (2024)
by: Bandini, Elena, et al.
Published: (2024)
Optimal strategies for the growth of dual-seeded lattice structures
by: de Jongh, Maike C., et al.
Published: (2025)
by: de Jongh, Maike C., et al.
Published: (2025)
Markov control of continuous time Markov processes with long run functionals by time discretization
by: Stettner, Lukasz
Published: (2025)
by: Stettner, Lukasz
Published: (2025)
Non-Expansive Mappings in Two-Time-Scale Stochastic Approximation: Finite-Time Analysis
by: Chandak, Siddharth
Published: (2025)
by: Chandak, Siddharth
Published: (2025)
On the Rate of Gaussian Approximation for Linear Regression Problems
by: Khusainov, Marat, et al.
Published: (2025)
by: Khusainov, Marat, et al.
Published: (2025)
Discrete-Time Approximations of Controlled Diffusions with Infinite Horizon Discounted and Average Cost
by: Pradhan, Somnath, et al.
Published: (2025)
by: Pradhan, Somnath, et al.
Published: (2025)
Infinite time horizon stochastic recursive control problems with jumps: dynamic programming and stochastic verification theorems
by: Luo, Sheng, et al.
Published: (2024)
by: Luo, Sheng, et al.
Published: (2024)
Policy Gradient for Continuous-Time Mean-Field Control
by: Bayraktar, Erhan, et al.
Published: (2026)
by: Bayraktar, Erhan, et al.
Published: (2026)
An Optimal-Control Approach to Infinite-Horizon Restless Bandits: Achieving Asymptotic Optimality with Minimal Assumptions
by: YAN, Chen
Published: (2024)
by: YAN, Chen
Published: (2024)
Stochastic Optimal Control of an Epidemic Under Partial Information
by: Njiasse, Ibrahim Mbouandi, et al.
Published: (2025)
by: Njiasse, Ibrahim Mbouandi, et al.
Published: (2025)
A geometric ensemble method for Bayesian inference
by: Popov, Andrey A
Published: (2025)
by: Popov, Andrey A
Published: (2025)
Entropic Risk-Averse Generalized Momentum Methods
by: Can, Bugra, et al.
Published: (2022)
by: Can, Bugra, et al.
Published: (2022)
High risk aversion Merton's problem without transversality conditions
by: Biffis, Enrico, et al.
Published: (2025)
by: Biffis, Enrico, et al.
Published: (2025)
Stability of long run functionals with respect to stationary Markov controls
by: Stettner, Lukasz
Published: (2024)
by: Stettner, Lukasz
Published: (2024)
Dynamic Programming Principle for Stochastic Control Problems on Riemannian Manifolds
by: Gao, Dingqian, et al.
Published: (2025)
by: Gao, Dingqian, et al.
Published: (2025)
Weakly-Coupled Multi-Action Restless Bandits -- Exponential Convergence in Probability
by: Fu, Jing, et al.
Published: (2026)
by: Fu, Jing, et al.
Published: (2026)
Reinforcement Learning, Optimal Control, and Bayesian Filtering in Data Assimilation
by: Hammoud, Abed
Published: (2026)
by: Hammoud, Abed
Published: (2026)
A Short Survey of Averaging Techniques in Stochastic Gradient Methods
by: Lakshmanan, K.
Published: (2026)
by: Lakshmanan, K.
Published: (2026)
Multi-Robot Relative Pose Estimation in SE(2) with Observability Analysis: A Comparison of Extended Kalman Filtering and Robust Pose Graph Optimization
by: Shin, Kihoon, et al.
Published: (2024)
by: Shin, Kihoon, et al.
Published: (2024)
Finite-Time Analysis of Projected Two-Time-Scale Stochastic Approximation
by: Bai, Yitao, et al.
Published: (2026)
by: Bai, Yitao, et al.
Published: (2026)
Sample Average Approximation for Stochastic Programming with Equality Constraints
by: Lew, Thomas, et al.
Published: (2022)
by: Lew, Thomas, et al.
Published: (2022)
Mean field social optimization: feedback person-by-person optimality and the dynamic programming equation
by: Huang, Minyi, et al.
Published: (2025)
by: Huang, Minyi, et al.
Published: (2025)
Optimal Control of a Stochastic Power System -- Algorithms and Mathematical Analysis
by: Wang, Zhen, et al.
Published: (2024)
by: Wang, Zhen, et al.
Published: (2024)
Optimal Control of Unbounded Functional Stochastic Evolution Systems in Hilbert Spaces: Second-Order Path-dependent HJB Equation
by: Tang, Shanjian, et al.
Published: (2024)
by: Tang, Shanjian, et al.
Published: (2024)
Convergence and turnpike properties of linear-quadratic mean field control problems with common noise
by: Bayraktar, Erhan, et al.
Published: (2026)
by: Bayraktar, Erhan, et al.
Published: (2026)
Similar Items
-
Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
by: Yu, Huizhen, et al.
Published: (2024) -
A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays
by: Yu, Huizhen, et al.
Published: (2023) -
Sample Complexity of Policy Gradient for Log-Growth Control
by: Pan, Qiuhua, et al.
Published: (2026) -
Average Cost Optimality of Partially Observed MDPS: Contraction of Non-linear Filters, Optimal Solutions and Approximations
by: Demirci, Yunus Emre, et al.
Published: (2023) -
On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies
by: Yu, Huizhen
Published: (2022)