On the Convergence of Monte Carlo UCB for Random-Length Episodic MDPs
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Zixuan, Wang, Che, Ross, Keith |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pre-training with Synthetic Data Helps Offline Reinforcement Learning
by: Wang, Zecheng, et al.
Published: (2023)
by: Wang, Zecheng, et al.
Published: (2023)
Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?
by: Omori, Yumi, et al.
Published: (2025)
by: Omori, Yumi, et al.
Published: (2025)
Is Optimal Transport Necessary for Inverse Reinforcement Learning?
by: Dong, Zixuan, et al.
Published: (2025)
by: Dong, Zixuan, et al.
Published: (2025)
The Prevalence of Neural Collapse in Neural Multivariate Regression
by: Andriopoulos, George, et al.
Published: (2024)
by: Andriopoulos, George, et al.
Published: (2024)
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)
by: Mondal, Washim Uddin, et al.
Published: (2024)
Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm
by: Xu, Yang, et al.
Published: (2025)
by: Xu, Yang, et al.
Published: (2025)
CAESAR: Enhancing Federated RL in Heterogeneous MDPs through Convergence-Aware Sampling with Screening
by: Mak, Hei Yi, et al.
Published: (2024)
by: Mak, Hei Yi, et al.
Published: (2024)
Monte Carlo Permutation Search
by: Cazenave, Tristan
Published: (2025)
by: Cazenave, Tristan
Published: (2025)
No-Regret Reinforcement Learning in Smooth MDPs
by: Maran, Davide, et al.
Published: (2024)
by: Maran, Davide, et al.
Published: (2024)
Linear $Q$-Learning Does Not Diverge in $L^2$: Convergence Rates to a Bounded Set
by: Liu, Xinyu, et al.
Published: (2025)
by: Liu, Xinyu, et al.
Published: (2025)
Epistemic Monte Carlo Tree Search
by: Oren, Yaniv, et al.
Published: (2022)
by: Oren, Yaniv, et al.
Published: (2022)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Efficient Solution and Learning of Robust Factored MDPs
by: Schnitzer, Yannik, et al.
Published: (2025)
by: Schnitzer, Yannik, et al.
Published: (2025)
Bourbaki: Self-Generated and Goal-Conditioned MDPs for Theorem Proving
by: Zimmer, Matthieu, et al.
Published: (2025)
by: Zimmer, Matthieu, et al.
Published: (2025)
Accelerating Monte-Carlo Tree Search with Optimized Posterior Policies
by: Frankston, Keith, et al.
Published: (2026)
by: Frankston, Keith, et al.
Published: (2026)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
by: Wang, Zige, et al.
Published: (2025)
by: Wang, Zige, et al.
Published: (2025)
SPO: Sequential Monte Carlo Policy Optimisation
by: Macfarlane, Matthew V, et al.
Published: (2024)
by: Macfarlane, Matthew V, et al.
Published: (2024)
Twice Sequential Monte Carlo for Tree Search
by: Oren, Yaniv, et al.
Published: (2025)
by: Oren, Yaniv, et al.
Published: (2025)
Monte Carlo Tree Search with Boltzmann Exploration
by: Painter, Michael, et al.
Published: (2024)
by: Painter, Michael, et al.
Published: (2024)
Doubly Robust Monte Carlo Tree Search
by: Liu, Manqing, et al.
Published: (2025)
by: Liu, Manqing, et al.
Published: (2025)
Missingness-MDPs: Bridging the Theory of Missing Data and POMDPs
by: Wendland, Joshua, et al.
Published: (2026)
by: Wendland, Joshua, et al.
Published: (2026)
Beyond Scalar Rewards: An Axiomatic Framework for Lexicographic MDPs
by: Shakerinava, Mehran, et al.
Published: (2025)
by: Shakerinava, Mehran, et al.
Published: (2025)
Geometry of Drifting MDPs with Path-Integral Stability Certificates
by: Zhang, Zuyuan, et al.
Published: (2026)
by: Zhang, Zuyuan, et al.
Published: (2026)
A UCB Bandit Algorithm for General ML-Based Estimators
by: Liu, Yajing, et al.
Published: (2026)
by: Liu, Yajing, et al.
Published: (2026)
Global Convergence of Multiplicative Updates for the Matrix Mechanism: A Collaborative Proof with Gemini 3
by: Rush, Keith
Published: (2026)
by: Rush, Keith
Published: (2026)
Episodic Novelty Through Temporal Distance
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
Convergence and sample complexity of natural policy gradient primal-dual methods for constrained MDPs
by: Ding, Dongsheng, et al.
Published: (2022)
by: Ding, Dongsheng, et al.
Published: (2022)
Monte Carlo Tree Diffusion for System 2 Planning
by: Yoon, Jaesik, et al.
Published: (2025)
by: Yoon, Jaesik, et al.
Published: (2025)
Anytime Sequential Halving in Monte-Carlo Tree Search
by: Sagers, Dominic, et al.
Published: (2024)
by: Sagers, Dominic, et al.
Published: (2024)
Monte Carlo Tree Search in the Presence of Transition Uncertainty
by: Kohankhaki, Farnaz, et al.
Published: (2023)
by: Kohankhaki, Farnaz, et al.
Published: (2023)
Improving GFlowNets with Monte Carlo Tree Search
by: Morozov, Nikita, et al.
Published: (2024)
by: Morozov, Nikita, et al.
Published: (2024)
Sequential Monte Carlo for Policy Optimization in Continuous POMDPs
by: Abdulsamad, Hany, et al.
Published: (2025)
by: Abdulsamad, Hany, et al.
Published: (2025)
Continuous Monte Carlo Graph Search
by: Kujanpää, Kalle, et al.
Published: (2022)
by: Kujanpää, Kalle, et al.
Published: (2022)
Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks
by: Ge, Luise, et al.
Published: (2025)
by: Ge, Luise, et al.
Published: (2025)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Exploration Implies Data Augmentation: Reachability and Generalisation in Contextual MDPs
by: Weltevrede, Max, et al.
Published: (2024)
by: Weltevrede, Max, et al.
Published: (2024)
Solving Multi-Model MDPs by Coordinate Ascent and Dynamic Programming
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
Risk-averse Total-reward MDPs with ERM and EVaR
by: Su, Xihong, et al.
Published: (2024)
by: Su, Xihong, et al.
Published: (2024)
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms
by: Shah, Anvay, et al.
Published: (2026)
by: Shah, Anvay, et al.
Published: (2026)
Monte Carlo Tree Search based Space Transfer for Black-box Optimization
by: Wang, Shukuan, et al.
Published: (2024)
by: Wang, Shukuan, et al.
Published: (2024)
Similar Items
-
Pre-training with Synthetic Data Helps Offline Reinforcement Learning
by: Wang, Zecheng, et al.
Published: (2023) -
Should We Ever Prefer Decision Transformer for Offline Reinforcement Learning?
by: Omori, Yumi, et al.
Published: (2025) -
Is Optimal Transport Necessary for Inverse Reinforcement Learning?
by: Dong, Zixuan, et al.
Published: (2025) -
The Prevalence of Neural Collapse in Neural Multivariate Regression
by: Andriopoulos, George, et al.
Published: (2024) -
Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
by: Mondal, Washim Uddin, et al.
Published: (2024)