Q-Measure-Learning for Continuous State RL: Efficient Implementation and Convergence
Fuente:
arXiv
Guardado en:
| Autor principal: | Wang, Shengbo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
por: Wang, Shengbo, et al.
Publicado: (2026)
por: Wang, Shengbo, et al.
Publicado: (2026)
Sample Complexity of Variance-reduced Distributionally Robust Q-learning
por: Wang, Shengbo, et al.
Publicado: (2023)
por: Wang, Shengbo, et al.
Publicado: (2023)
Convergence and stability of Q-learning in Hierarchical Reinforcement Learning
por: Manenti, Massimiliano, et al.
Publicado: (2025)
por: Manenti, Massimiliano, et al.
Publicado: (2025)
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
por: Chen, Zijun, et al.
Publicado: (2025)
por: Chen, Zijun, et al.
Publicado: (2025)
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
por: Wang, Shengbo, et al.
Publicado: (2026)
por: Wang, Shengbo, et al.
Publicado: (2026)
Bellman Optimality of Average-Reward Robust Markov Decision Processes with a Constant Gain
por: Wang, Shengbo, et al.
Publicado: (2025)
por: Wang, Shengbo, et al.
Publicado: (2025)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
por: Wan, Yi, et al.
Publicado: (2024)
por: Wan, Yi, et al.
Publicado: (2024)
Communication Efficient Federated Learning with Linear Convergence on Heterogeneous Data
por: Liu, Jie, et al.
Publicado: (2025)
por: Liu, Jie, et al.
Publicado: (2025)
On the Foundation of Distributionally Robust Reinforcement Learning
por: Wang, Shengbo, et al.
Publicado: (2023)
por: Wang, Shengbo, et al.
Publicado: (2023)
Last Iterate Convergence of Incremental Methods and Applications in Continual Learning
por: Cai, Xufeng, et al.
Publicado: (2024)
por: Cai, Xufeng, et al.
Publicado: (2024)
Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control
por: Hua, Chengxiu, et al.
Publicado: (2025)
por: Hua, Chengxiu, et al.
Publicado: (2025)
State Estimation Using Particle Filtering in Adaptive Machine Learning Methods: Integrating Q-Learning and NEAT Algorithms with Noisy Radar Measurements
por: Song, Wonjin, et al.
Publicado: (2025)
por: Song, Wonjin, et al.
Publicado: (2025)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
por: Zhang, Yixuan, et al.
Publicado: (2024)
por: Zhang, Yixuan, et al.
Publicado: (2024)
Optimal Sample Complexity for Average Reward Markov Decision Processes
por: Wang, Shengbo, et al.
Publicado: (2023)
por: Wang, Shengbo, et al.
Publicado: (2023)
Risk-Sensitive Q-Learning in Continuous Time with Application to Dynamic Portfolio Selection
por: Xie, Chuhan
Publicado: (2025)
por: Xie, Chuhan
Publicado: (2025)
Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation
por: Kong, Deyi, et al.
Publicado: (2026)
por: Kong, Deyi, et al.
Publicado: (2026)
ADMM Algorithms for Residual Network Training: Convergence Analysis and Parallel Implementation
por: Xu, Jintao, et al.
Publicado: (2023)
por: Xu, Jintao, et al.
Publicado: (2023)
Performance of NPG in Countable State-Space Average-Cost RL
por: Murthy, Yashaswini, et al.
Publicado: (2024)
por: Murthy, Yashaswini, et al.
Publicado: (2024)
Revisiting Subgradient Method: Complexity and Convergence Beyond Lipschitz Continuity
por: Li, Xiao, et al.
Publicado: (2023)
por: Li, Xiao, et al.
Publicado: (2023)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
por: Jung, Hyunji, et al.
Publicado: (2025)
por: Jung, Hyunji, et al.
Publicado: (2025)
SINDy-RL: Interpretable and Efficient Model-Based Reinforcement Learning
por: Zolman, Nicholas, et al.
Publicado: (2024)
por: Zolman, Nicholas, et al.
Publicado: (2024)
Provably Convergent Federated Trilevel Learning
por: Jiao, Yang, et al.
Publicado: (2023)
por: Jiao, Yang, et al.
Publicado: (2023)
Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field Theory
por: Zhang, Yufeng, et al.
Publicado: (2020)
por: Zhang, Yufeng, et al.
Publicado: (2020)
Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
por: Fouque, Jean-Pierre, et al.
Publicado: (2025)
por: Fouque, Jean-Pierre, et al.
Publicado: (2025)
Efficient Sign-Based Optimization: Accelerating Convergence via Variance Reduction
por: Jiang, Wei, et al.
Publicado: (2024)
por: Jiang, Wei, et al.
Publicado: (2024)
VAMO: Efficient Zeroth-Order Variance Reduction for SGD with Faster Convergence
por: Chen, Jiahe, et al.
Publicado: (2025)
por: Chen, Jiahe, et al.
Publicado: (2025)
Learning Provably Improves the Convergence of Gradient Descent
por: Song, Qingyu, et al.
Publicado: (2025)
por: Song, Qingyu, et al.
Publicado: (2025)
Memory-Reduced Meta-Learning with Guaranteed Convergence
por: Yang, Honglin, et al.
Publicado: (2024)
por: Yang, Honglin, et al.
Publicado: (2024)
The Role of Target Update Frequencies in Q-Learning
por: Weissmann, Simon, et al.
Publicado: (2026)
por: Weissmann, Simon, et al.
Publicado: (2026)
ADDQ: Adaptive Distributional Double Q-Learning
por: Döring, Leif, et al.
Publicado: (2025)
por: Döring, Leif, et al.
Publicado: (2025)
Pointer Networks with Q-Learning for Combinatorial Optimization
por: Barro, Alessandro
Publicado: (2023)
por: Barro, Alessandro
Publicado: (2023)
On the Global Convergence of Risk-Averse Natural Policy Gradient Methods with Expected Conditional Risk Measures
por: Yu, Xian, et al.
Publicado: (2023)
por: Yu, Xian, et al.
Publicado: (2023)
Robust Q-Learning under Corrupted Rewards
por: Maity, Sreejeet, et al.
Publicado: (2024)
por: Maity, Sreejeet, et al.
Publicado: (2024)
Provably Efficient Representation Selection in Low-rank Markov Decision Processes: From Online to Offline RL
por: Zhang, Weitong, et al.
Publicado: (2021)
por: Zhang, Weitong, et al.
Publicado: (2021)
Learning Over-Relaxation Policies for ADMM with Convergence Guarantees
por: Lin, Junan, et al.
Publicado: (2026)
por: Lin, Junan, et al.
Publicado: (2026)
On the Convergence of Gradient Descent on Learning Transformers with Residual Connections
por: Qin, Zhen, et al.
Publicado: (2025)
por: Qin, Zhen, et al.
Publicado: (2025)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
por: Lin, Yifan, et al.
Publicado: (2024)
por: Lin, Yifan, et al.
Publicado: (2024)
Central Limit Theorems for Asynchronous Averaged Q-Learning
por: Liu, Xingtu
Publicado: (2025)
por: Liu, Xingtu
Publicado: (2025)
Convergence Rate in Nonlinear Two-Time-Scale Stochastic Approximation with State (Time)-Dependence
por: Chen, Zixi, et al.
Publicado: (2025)
por: Chen, Zixi, et al.
Publicado: (2025)
Efficient Continual Finite-Sum Minimization
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2024)
por: Mavrothalassitis, Ioannis, et al.
Publicado: (2024)
Ejemplares similares
-
Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
por: Wang, Shengbo, et al.
Publicado: (2026) -
Sample Complexity of Variance-reduced Distributionally Robust Q-learning
por: Wang, Shengbo, et al.
Publicado: (2023) -
Convergence and stability of Q-learning in Hierarchical Reinforcement Learning
por: Manenti, Massimiliano, et al.
Publicado: (2025) -
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
por: Chen, Zijun, et al.
Publicado: (2025) -
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
por: Wang, Shengbo, et al.
Publicado: (2026)