Asynchronous Stochastic Approximation with Applications to Average-Reward Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Huizhen, Wan, Yi, Sutton, Richard S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
by: Yu, Huizhen, et al.
Published: (2025)
by: Yu, Huizhen, et al.
Published: (2025)
A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays
by: Yu, Huizhen, et al.
Published: (2023)
by: Yu, Huizhen, et al.
Published: (2023)
On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies
by: Yu, Huizhen
Published: (2022)
by: Yu, Huizhen
Published: (2022)
Average Cost Optimality of Partially Observed MDPS: Contraction of Non-linear Filters, Optimal Solutions and Approximations
by: Demirci, Yunus Emre, et al.
Published: (2023)
by: Demirci, Yunus Emre, et al.
Published: (2023)
Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation
by: Zheng, Yaowei, et al.
Published: (2026)
by: Zheng, Yaowei, et al.
Published: (2026)
Sample Complexity of Policy Gradient for Log-Growth Control
by: Pan, Qiuhua, et al.
Published: (2026)
by: Pan, Qiuhua, et al.
Published: (2026)
Non-Expansive Mappings in Two-Time-Scale Stochastic Approximation: Finite-Time Analysis
by: Chandak, Siddharth
Published: (2025)
by: Chandak, Siddharth
Published: (2025)
Optimistic Training and Convergence of Q-Learning -- Extended Version
by: Mehta, Prashant, et al.
Published: (2026)
by: Mehta, Prashant, et al.
Published: (2026)
Stability and Sensitivity Analysis of Relative Temporal-Difference Learning: Extended Version
by: Sakha, Masoud S., et al.
Published: (2026)
by: Sakha, Masoud S., et al.
Published: (2026)
Stochastic Control with Signatures
by: Bank, P., et al.
Published: (2024)
by: Bank, P., et al.
Published: (2024)
A Sequential Testing Problem with Signal Control
by: Campbell, Steven, et al.
Published: (2025)
by: Campbell, Steven, et al.
Published: (2025)
Robust Ergodic Control of Jump-Diffusion Systems under Drift and Intensity Uncertainty
by: Azze, Abel, et al.
Published: (2026)
by: Azze, Abel, et al.
Published: (2026)
Markovian Foundations for Quasi-Stochastic Approximation with Applications to Extremum Seeking Control
by: Lauand, Caio Kalil, et al.
Published: (2022)
by: Lauand, Caio Kalil, et al.
Published: (2022)
Sample Average Approximation for Stochastic Programming with Equality Constraints
by: Lew, Thomas, et al.
Published: (2022)
by: Lew, Thomas, et al.
Published: (2022)
Gaussian Approximation and Multiplier Bootstrap for Polyak-Ruppert Averaged Linear Stochastic Approximation with Applications to TD Learning
by: Samsonov, Sergey, et al.
Published: (2024)
by: Samsonov, Sergey, et al.
Published: (2024)
Optimal strategies for the growth of dual-seeded lattice structures
by: de Jongh, Maike C., et al.
Published: (2025)
by: de Jongh, Maike C., et al.
Published: (2025)
On the Rate of Gaussian Approximation for Linear Regression Problems
by: Khusainov, Marat, et al.
Published: (2025)
by: Khusainov, Marat, et al.
Published: (2025)
High Probability Bounds for Stochastic Subgradient Schemes with Heavy Tailed Noise
by: Parletta, Daniela A., et al.
Published: (2022)
by: Parletta, Daniela A., et al.
Published: (2022)
Finite-Time Analysis of Projected Two-Time-Scale Stochastic Approximation
by: Bai, Yitao, et al.
Published: (2026)
by: Bai, Yitao, et al.
Published: (2026)
Reinforcement Learning Methods for the Stochastic Optimal Control of an Industrial Power-to-Heat System
by: Pilling, Eric, et al.
Published: (2024)
by: Pilling, Eric, et al.
Published: (2024)
Cost-optimal Management of a Residential Heating System With a Geothermal Energy Storage Under Uncertainty
by: Takam, Paul Honore, et al.
Published: (2025)
by: Takam, Paul Honore, et al.
Published: (2025)
Stochastic Optimal Control of an Epidemic Under Partial Information
by: Njiasse, Ibrahim Mbouandi, et al.
Published: (2025)
by: Njiasse, Ibrahim Mbouandi, et al.
Published: (2025)
Optimal Control of a Stochastic Power System -- Algorithms and Mathematical Analysis
by: Wang, Zhen, et al.
Published: (2024)
by: Wang, Zhen, et al.
Published: (2024)
Entropic Risk-Averse Generalized Momentum Methods
by: Can, Bugra, et al.
Published: (2022)
by: Can, Bugra, et al.
Published: (2022)
Causal Hamilton-Jacobi-Bellman Equations for Anticipative Stochastic Optimal Control
by: Bank, Peter, et al.
Published: (2025)
by: Bank, Peter, et al.
Published: (2025)
Long run control of nonhomogeneous Markov processes
by: Stettner, Łukasz
Published: (2025)
by: Stettner, Łukasz
Published: (2025)
Drift Optimization of Regulated Stochastic Models Using Sample Average Approximation
by: Zhou, Zihe, et al.
Published: (2025)
by: Zhou, Zihe, et al.
Published: (2025)
Gaussian Approximation and Multiplier Bootstrap for Stochastic Gradient Descent
by: Sheshukova, Marina, et al.
Published: (2025)
by: Sheshukova, Marina, et al.
Published: (2025)
Optimal Control of Unbounded Functional Stochastic Evolution Systems in Hilbert Spaces: Second-Order Path-dependent HJB Equation
by: Tang, Shanjian, et al.
Published: (2024)
by: Tang, Shanjian, et al.
Published: (2024)
An Optimal-Control Approach to Infinite-Horizon Restless Bandits: Achieving Asymptotic Optimality with Minimal Assumptions
by: YAN, Chen
Published: (2024)
by: YAN, Chen
Published: (2024)
Revisiting Stochastic Approximation and Stochastic Gradient Descent
by: Karandikar, Rajeeva Laxman, et al.
Published: (2025)
by: Karandikar, Rajeeva Laxman, et al.
Published: (2025)
Thompson Sampling for Infinite-Horizon Discounted Decision Processes
by: Adelman, Daniel, et al.
Published: (2024)
by: Adelman, Daniel, et al.
Published: (2024)
Non-local Hamilton-Jacobi-Bellman equations for the stochastic optimal control of path-dependent piecewise deterministic processes
by: Bandini, Elena, et al.
Published: (2024)
by: Bandini, Elena, et al.
Published: (2024)
Deep Relaxation of Controlled Stochastic Gradient Descent via Singular Perturbations
by: Bardi, Martino, et al.
Published: (2022)
by: Bardi, Martino, et al.
Published: (2022)
Unichain and Aperiodicity are Sufficient for Asymptotic Optimality of Average-Reward Restless Bandits
by: Hong, Yige, et al.
Published: (2024)
by: Hong, Yige, et al.
Published: (2024)
A Fisher-Rao gradient flow for entropy-regularised Markov decision processes in Polish spaces
by: Kerimkulov, Bekzhan, et al.
Published: (2023)
by: Kerimkulov, Bekzhan, et al.
Published: (2023)
Statistical inference for Linear Stochastic Approximation with Markovian Noise
by: Samsonov, Sergey, et al.
Published: (2025)
by: Samsonov, Sergey, et al.
Published: (2025)
Convergence Rates for Stochastic Approximation: Biased Noise with Unbounded Variance, and Applications
by: Karandikar, Rajeeva L., et al.
Published: (2023)
by: Karandikar, Rajeeva L., et al.
Published: (2023)
Discrete-Time Approximations of Controlled Diffusions with Infinite Horizon Discounted and Average Cost
by: Pradhan, Somnath, et al.
Published: (2025)
by: Pradhan, Somnath, et al.
Published: (2025)
Improved Central Limit Theorem and Bootstrap Approximations for Linear Stochastic Approximation
by: Butyrin, Bogdan, et al.
Published: (2025)
by: Butyrin, Bogdan, et al.
Published: (2025)
Similar Items
-
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
by: Yu, Huizhen, et al.
Published: (2025) -
A Note on Stability in Asynchronous Stochastic Approximation without Communication Delays
by: Yu, Huizhen, et al.
Published: (2023) -
On Strategic Measures and Optimality Properties in Discrete-Time Stochastic Control with Universally Measurable Policies
by: Yu, Huizhen
Published: (2022) -
Average Cost Optimality of Partially Observed MDPS: Contraction of Non-linear Filters, Optimal Solutions and Approximations
by: Demirci, Yunus Emre, et al.
Published: (2023) -
Beyond Bellman: High-Order Generator Regression for Continuous-Time Policy Evaluation
by: Zheng, Yaowei, et al.
Published: (2026)