Logarithmic regret bounds for continuous-time average-reward Markov decision processes
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Xuefeng, Zhou, Xun Yu |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stochastic first-order methods for average-reward Markov decision processes
by: Li, Tianjiao, et al.
Published: (2022)
by: Li, Tianjiao, et al.
Published: (2022)
A safe exploration approach to constrained Markov decision processes
by: Ni, Tingting, et al.
Published: (2023)
by: Ni, Tingting, et al.
Published: (2023)
Reward-Directed Score-Based Diffusion Models via q-Learning
by: Gao, Xuefeng, et al.
Published: (2024)
by: Gao, Xuefeng, et al.
Published: (2024)
Reinforcement Learning for Jump-Diffusions, with Financial Applications
by: Gao, Xuefeng, et al.
Published: (2024)
by: Gao, Xuefeng, et al.
Published: (2024)
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
by: Yu, Huizhen, et al.
Published: (2025)
by: Yu, Huizhen, et al.
Published: (2025)
Data-driven robust Markov decision processes on Borel spaces: performance guarantees via an axiomatic approach
by: Ramani, Sivaramakrishnan
Published: (2026)
by: Ramani, Sivaramakrishnan
Published: (2026)
Orlicz regrets to consistently bound statistics of random variables with an application to environmental indicators
by: Yoshioka, Hidekazu, et al.
Published: (2023)
by: Yoshioka, Hidekazu, et al.
Published: (2023)
Localized exploration in contextual dynamic pricing achieves dimension-free regret
by: Chai, Jinhang, et al.
Published: (2024)
by: Chai, Jinhang, et al.
Published: (2024)
Logarithmic-time Schedules for Scaling Language Models with Momentum
by: Ferbach, Damien, et al.
Published: (2026)
by: Ferbach, Damien, et al.
Published: (2026)
Online Convex Optimization and Integral Quadratic Constraints: An automated approach to regret analysis
by: Jakob, Fabian, et al.
Published: (2025)
by: Jakob, Fabian, et al.
Published: (2025)
Flipping-based Policy for Chance-Constrained Markov Decision Processes
by: Shen, Xun, et al.
Published: (2024)
by: Shen, Xun, et al.
Published: (2024)
Randomized algorithms and PAC bounds for inverse reinforcement learning in continuous spaces
by: Kamoutsi, Angeliki, et al.
Published: (2024)
by: Kamoutsi, Angeliki, et al.
Published: (2024)
Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management
by: Meng, Huiling, et al.
Published: (2024)
by: Meng, Huiling, et al.
Published: (2024)
Regret Bounds for Episodic Risk-Sensitive Linear Quadratic Regulator
by: Xu, Wenhao, et al.
Published: (2024)
by: Xu, Wenhao, et al.
Published: (2024)
Logarithmic Regret for Unconstrained Submodular Maximization Stochastic Bandit
by: Zhou, Julien, et al.
Published: (2024)
by: Zhou, Julien, et al.
Published: (2024)
Regret of exploratory policy improvement and $q$-learning
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
Achieving Logarithmic Regret in KL-Regularized Zero-Sum Markov Games
by: Nayak, Anupam, et al.
Published: (2025)
by: Nayak, Anupam, et al.
Published: (2025)
Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations
by: Ren, Zhenjie, et al.
Published: (2026)
by: Ren, Zhenjie, et al.
Published: (2026)
Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms
by: Ren, Zhenjie, et al.
Published: (2026)
by: Ren, Zhenjie, et al.
Published: (2026)
Logarithmic-Regret Quantum Learning Algorithms for Zero-Sum Games
by: Gao, Minbo, et al.
Published: (2023)
by: Gao, Minbo, et al.
Published: (2023)
Data-Driven Exploration for a Class of Continuous-Time Indefinite Linear--Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2025)
by: Huang, Yilie, et al.
Published: (2025)
Online Nonstochastic Prediction: Logarithmic Regret via Predictive Online Least Squares
by: Pai, Chih-Fan, et al.
Published: (2026)
by: Pai, Chih-Fan, et al.
Published: (2026)
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
by: Lugosi, Gabor, et al.
Published: (2024)
by: Lugosi, Gabor, et al.
Published: (2024)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
by: Wan, Yi, et al.
Published: (2024)
by: Wan, Yi, et al.
Published: (2024)
Leveraging Hamilton-Jacobi PDEs with time-dependent Hamiltonians for continual scientific machine learning
by: Chen, Paula, et al.
Published: (2023)
by: Chen, Paula, et al.
Published: (2023)
Sublinear Regret for a Class of Continuous-Time Linear-Quadratic Reinforcement Learning Problems
by: Huang, Yilie, et al.
Published: (2024)
by: Huang, Yilie, et al.
Published: (2024)
Adam with model exponential moving average is effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
by: Zhang, Weitong, et al.
Published: (2021)
by: Zhang, Weitong, et al.
Published: (2021)
Soft decision trees for survival analysis
by: Consolo, Antonio, et al.
Published: (2025)
by: Consolo, Antonio, et al.
Published: (2025)
MARS: Unleashing the Power of Variance Reduction for Training Large Models
by: Yuan, Huizhuo, et al.
Published: (2024)
by: Yuan, Huizhuo, et al.
Published: (2024)
A note on continuous-time online learning
by: Ying, Lexing
Published: (2024)
by: Ying, Lexing
Published: (2024)
Weakly Time-Coupled Approximation of Markov Decision Processes
by: Soheili, Negar, et al.
Published: (2026)
by: Soheili, Negar, et al.
Published: (2026)
Online Markov Decision Processes with Terminal Law Constraints
by: Moreno, Bianca Marin, et al.
Published: (2026)
by: Moreno, Bianca Marin, et al.
Published: (2026)
Reinforcement Learning with Function Approximation for Non-Markov Processes
by: Kara, Ali Devran
Published: (2026)
by: Kara, Ali Devran
Published: (2026)
Unified continuous-time q-learning for mean-field game and mean-field control problems
by: Wei, Xiaoli, et al.
Published: (2024)
by: Wei, Xiaoli, et al.
Published: (2024)
Adaptive decision-making for stochastic service network design
by: Durán-Micco, Javier, et al.
Published: (2026)
by: Durán-Micco, Javier, et al.
Published: (2026)
Optimal Sample Complexity for Average Reward Markov Decision Processes
by: Wang, Shengbo, et al.
Published: (2023)
by: Wang, Shengbo, et al.
Published: (2023)
Beyond IID: data-driven decision-making in heterogeneous environments
by: Besbes, Omar, et al.
Published: (2022)
by: Besbes, Omar, et al.
Published: (2022)
Full error analysis of policy gradient learning algorithms for exploratory linear quadratic mean-field control problem in continuous time with common noise
by: Frikha, Noufel, et al.
Published: (2024)
by: Frikha, Noufel, et al.
Published: (2024)
Towards Simple and Provable Parameter-Free Adaptive Gradient Methods
by: Tao, Yuanzhe, et al.
Published: (2024)
by: Tao, Yuanzhe, et al.
Published: (2024)
Similar Items
-
Stochastic first-order methods for average-reward Markov decision processes
by: Li, Tianjiao, et al.
Published: (2022) -
A safe exploration approach to constrained Markov decision processes
by: Ni, Tingting, et al.
Published: (2023) -
Reward-Directed Score-Based Diffusion Models via q-Learning
by: Gao, Xuefeng, et al.
Published: (2024) -
Reinforcement Learning for Jump-Diffusions, with Financial Applications
by: Gao, Xuefeng, et al.
Published: (2024) -
Average-reward reinforcement learning in semi-Markov decision processes via relative value iteration
by: Yu, Huizhen, et al.
Published: (2025)