Decoupling Exploration and Policy Optimization: Uncertainty Guided Tree Search for Hard Exploration
Fuente:
arXiv
Saved in:
| Main Authors: | Mhammedi, Zakaria, Cohan, James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Online Convex Optimization with a Separation Oracle
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
Efficient Model-Free Exploration in Low-Rank MDPs
by: Mhammedi, Zakaria, et al.
Published: (2023)
by: Mhammedi, Zakaria, et al.
Published: (2023)
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024)
by: Mhammedi, Zakaria
Published: (2024)
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
by: Foster, Dylan J., et al.
Published: (2025)
by: Foster, Dylan J., et al.
Published: (2025)
Fully Unconstrained Online Learning
by: Cutkosky, Ashok, et al.
Published: (2024)
by: Cutkosky, Ashok, et al.
Published: (2024)
End-to-End Efficient RL for Linear Bellman Complete MDPs with Deterministic Transitions
by: Mhammedi, Zakaria, et al.
Published: (2026)
by: Mhammedi, Zakaria, et al.
Published: (2026)
Adaptive Matrix Online Learning through Smoothing with Guarantees for Nonsmooth Nonconvex Optimization
by: Jiang, Ruichen, et al.
Published: (2026)
by: Jiang, Ruichen, et al.
Published: (2026)
The Power of Resets in Online Reinforcement Learning
by: Mhammedi, Zakaria, et al.
Published: (2024)
by: Mhammedi, Zakaria, et al.
Published: (2024)
Beating Adversarial Low-Rank MDPs with Unknown Transition and Bandit Feedback
by: Liu, Haolin, et al.
Published: (2024)
by: Liu, Haolin, et al.
Published: (2024)
Proximal Policy Optimization with Adaptive Exploration
by: Lixandru, Andrei
Published: (2024)
by: Lixandru, Andrei
Published: (2024)
Provably Efficient Exploration in Policy Optimization
by: Cai, Qi, et al.
Published: (2019)
by: Cai, Qi, et al.
Published: (2019)
NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search
by: Tang, Sizhe, et al.
Published: (2026)
by: Tang, Sizhe, et al.
Published: (2026)
Optimization of Epsilon-Greedy Exploration
by: Che, Ethan, et al.
Published: (2025)
by: Che, Ethan, et al.
Published: (2025)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
by: Qu, Yuxiao, et al.
Published: (2026)
by: Qu, Yuxiao, et al.
Published: (2026)
Efficient Reinforcement Learning via Decoupling Exploration and Utilization
by: Yang, Jingpu, et al.
Published: (2023)
by: Yang, Jingpu, et al.
Published: (2023)
Monte Carlo Tree Search with Boltzmann Exploration
by: Painter, Michael, et al.
Published: (2024)
by: Painter, Michael, et al.
Published: (2024)
How Log-Barrier Helps Exploration in Policy Optimization
by: Cesani, Leonardo, et al.
Published: (2026)
by: Cesani, Leonardo, et al.
Published: (2026)
Can Tabular Foundation Models Guide Exploration in Robot Policy Learning?
by: Ou, Buqing, et al.
Published: (2026)
by: Ou, Buqing, et al.
Published: (2026)
Uncertainty-Guided Likelihood Tree Search
by: Grosse, Julia, et al.
Published: (2024)
by: Grosse, Julia, et al.
Published: (2024)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Search Inspired Exploration in Reinforcement Learning
by: Sotirchos, Georgios, et al.
Published: (2026)
by: Sotirchos, Georgios, et al.
Published: (2026)
Distribution-Centric Policy Optimization Dominates Exploration-Exploitation Trade-off
by: Li, Zhaochun, et al.
Published: (2026)
by: Li, Zhaochun, et al.
Published: (2026)
Behind the Myth of Exploration in Policy Gradients
by: Bolland, Adrien, et al.
Published: (2024)
by: Bolland, Adrien, et al.
Published: (2024)
Guardian: Decoupling Exploration from Safety in Reinforcement Learning
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
Exploration Behavior of Untrained Policies
by: Adamczyk, Jacob
Published: (2025)
by: Adamczyk, Jacob
Published: (2025)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
by: Bartoldson, Brian, et al.
Published: (2025)
by: Bartoldson, Brian, et al.
Published: (2025)
Exploring Exploration in Bayesian Optimization
by: Papenmeier, Leonard, et al.
Published: (2025)
by: Papenmeier, Leonard, et al.
Published: (2025)
Decoupling Exploration and Exploitation for Unsupervised Pre-training with Successor Features
by: Kim, JaeYoon, et al.
Published: (2024)
by: Kim, JaeYoon, et al.
Published: (2024)
Reinforcement Learning under Latent Dynamics: Toward Statistical and Algorithmic Modularity
by: Amortila, Philip, et al.
Published: (2024)
by: Amortila, Philip, et al.
Published: (2024)
UGCE: User-Guided Incremental Counterfactual Exploration
by: Fragkathoulas, Christos, et al.
Published: (2025)
by: Fragkathoulas, Christos, et al.
Published: (2025)
Preference-Guided Reinforcement Learning for Efficient Exploration
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Guided Exploration in Reinforcement Learning via Monte Carlo Critic Optimization
by: Kuznetsov, Igor
Published: (2022)
by: Kuznetsov, Igor
Published: (2022)
Reinforcement Learning by Guided Safe Exploration
by: Yang, Qisong, et al.
Published: (2023)
by: Yang, Qisong, et al.
Published: (2023)
Safe Policy Exploration Improvement via Subgoals
by: Angulo, Brian, et al.
Published: (2024)
by: Angulo, Brian, et al.
Published: (2024)
Feasible-First Exploration for Constrained ML Deployment Optimization in Crash-Prone Hierarchical Search Spaces
by: Lysenstøen, Christian
Published: (2026)
by: Lysenstøen, Christian
Published: (2026)
Optimal Control-Based Baseline for Guided Exploration in Policy Gradient Methods
by: Lyu, Xubo, et al.
Published: (2020)
by: Lyu, Xubo, et al.
Published: (2020)
DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off
by: Li, Xiaofan, et al.
Published: (2026)
by: Li, Xiaofan, et al.
Published: (2026)
Safe Exploration via Policy Priors
by: Wendl, Manuel, et al.
Published: (2026)
by: Wendl, Manuel, et al.
Published: (2026)
Expanding LLM Agent Boundaries with Strategy-Guided Exploration
by: Szot, Andrew, et al.
Published: (2026)
by: Szot, Andrew, et al.
Published: (2026)
Hierarchical Uncertainty Exploration via Feedforward Posterior Trees
by: Nehme, Elias, et al.
Published: (2024)
by: Nehme, Elias, et al.
Published: (2024)
Similar Items
-
Online Convex Optimization with a Separation Oracle
by: Mhammedi, Zakaria
Published: (2024) -
Efficient Model-Free Exploration in Low-Rank MDPs
by: Mhammedi, Zakaria, et al.
Published: (2023) -
Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
by: Mhammedi, Zakaria
Published: (2024) -
Is a Good Foundation Necessary for Efficient Reinforcement Learning? The Computational Role of the Base Model in Exploration
by: Foster, Dylan J., et al.
Published: (2025) -
Fully Unconstrained Online Learning
by: Cutkosky, Ashok, et al.
Published: (2024)