Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
Fuente:
arXiv
Saved in:
| Main Authors: | Koren, Uri, Kumar, Navdeep, Gadot, Uri, Ramponi, Giorgia, Levy, Kfir Yehuda, Mannor, Shie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2025)
by: Kumar, Navdeep, et al.
Published: (2025)
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
by: Wang, Kaixin, et al.
Published: (2023)
by: Wang, Kaixin, et al.
Published: (2023)
On the Convergence of Single-Timescale Actor-Critic
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
by: Gadot, Uri, et al.
Published: (2023)
by: Gadot, Uri, et al.
Published: (2023)
Simulus: Combining Improvements in Sample-Efficient World Model Agents
by: Cohen, Lior, et al.
Published: (2025)
by: Cohen, Lior, et al.
Published: (2025)
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum
by: Kumar, Navdeep, et al.
Published: (2026)
by: Kumar, Navdeep, et al.
Published: (2026)
Policy Optimized Text-to-Image Pipeline Design
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
Policy Gradient with Tree Expansion
by: Dalal, Gal, et al.
Published: (2023)
by: Dalal, Gal, et al.
Published: (2023)
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2024)
by: Kumar, Navdeep, et al.
Published: (2024)
Tree Search-Based Policy Optimization under Stochastic Execution Delay
by: Valensi, David, et al.
Published: (2024)
by: Valensi, David, et al.
Published: (2024)
RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression
by: Gadot, Uri, et al.
Published: (2025)
by: Gadot, Uri, et al.
Published: (2025)
MinMaxMin $Q$-learning
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Conservative DDPG -- Pessimistic RL without Ensemble
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Controlling False Discovery in Arbitrarily Structured Hypothesis Spaces via Reproducing Kernels
by: Perets, Binyamin, et al.
Published: (2026)
by: Perets, Binyamin, et al.
Published: (2026)
Representation-Driven Reinforcement Learning
by: Nabati, Ofir, et al.
Published: (2023)
by: Nabati, Ofir, et al.
Published: (2023)
Sobolev Space Regularised Pre Density Models
by: Kozdoba, Mark, et al.
Published: (2023)
by: Kozdoba, Mark, et al.
Published: (2023)
Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning
by: Macuglia, Maël, et al.
Published: (2025)
by: Macuglia, Maël, et al.
Published: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Turning Sand to Gold: Recycling Data to Bridge On-Policy and Off-Policy Learning via Causal Bound
by: Fiskus, Tal, et al.
Published: (2025)
by: Fiskus, Tal, et al.
Published: (2025)
State Entropy Regularization for Robust Reinforcement Learning
by: Ashlag, Yonatan, et al.
Published: (2025)
by: Ashlag, Yonatan, et al.
Published: (2025)
Improving Token-Based World Models with Parallel Observation Prediction
by: Cohen, Lior, et al.
Published: (2024)
by: Cohen, Lior, et al.
Published: (2024)
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
by: Du, Yihan, et al.
Published: (2024)
by: Du, Yihan, et al.
Published: (2024)
SQT -- std $Q$-target
by: Soffair, Nitsan, et al.
Published: (2024)
by: Soffair, Nitsan, et al.
Published: (2024)
Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models
by: Cohen, Lior, et al.
Published: (2026)
by: Cohen, Lior, et al.
Published: (2026)
Policy Mirror Descent with Lookahead
by: Protopapas, Kimon, et al.
Published: (2024)
by: Protopapas, Kimon, et al.
Published: (2024)
VLM-Guided Experience Replay
by: Sharony, Elad, et al.
Published: (2026)
by: Sharony, Elad, et al.
Published: (2026)
The Dimension Strikes Back with Gradients: Generalization of Gradient Methods in Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024)
by: Schliserman, Matan, et al.
Published: (2024)
Preference Elicitation for Offline Reinforcement Learning
by: Pace, Alizée, et al.
Published: (2024)
by: Pace, Alizée, et al.
Published: (2024)
Regret Minimization and Convergence to Equilibria in General-sum Markov Games
by: Erez, Liad, et al.
Published: (2022)
by: Erez, Liad, et al.
Published: (2022)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
Privacy-Preserving Federated Convex Optimization: Balancing Partial-Participation and Efficiency via Noise Cancellation
by: Reshef, Roie, et al.
Published: (2025)
by: Reshef, Roie, et al.
Published: (2025)
Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards
by: Wu, Shirley, et al.
Published: (2025)
by: Wu, Shirley, et al.
Published: (2025)
Lookahead Counterfactual Fairness
by: Zuo, Zhiqun, et al.
Published: (2024)
by: Zuo, Zhiqun, et al.
Published: (2024)
Learning Multiple Initial Solutions to Optimization Problems
by: Sharony, Elad, et al.
Published: (2024)
by: Sharony, Elad, et al.
Published: (2024)
Rate optimal learning of equilibria from data
by: Freihaut, Till, et al.
Published: (2025)
by: Freihaut, Till, et al.
Published: (2025)
Contextual Bilevel Reinforcement Learning for Incentive Alignment
by: Thoma, Vinzenz, et al.
Published: (2024)
by: Thoma, Vinzenz, et al.
Published: (2024)
Unsupervised Ensemble Learning Through Deep Energy-based Models
by: Maymon, Ariel, et al.
Published: (2026)
by: Maymon, Ariel, et al.
Published: (2026)
Learning with Foresight: Enhancing Neural Routing Policy via Multi-Node Lookahead Prediction
by: Jiang, Xia, et al.
Published: (2026)
by: Jiang, Xia, et al.
Published: (2026)
Improving Inverse Folding for Peptide Design with Diversity-regularized Direct Preference Optimization
by: Park, Ryan, et al.
Published: (2024)
by: Park, Ryan, et al.
Published: (2024)
Similar Items
-
Dual Formulation for Non-Rectangular Lp Robust Markov Decision Processes
by: Kumar, Navdeep, et al.
Published: (2025) -
Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel
by: Wang, Kaixin, et al.
Published: (2023) -
On the Convergence of Single-Timescale Actor-Critic
by: Kumar, Navdeep, et al.
Published: (2024) -
Solving Non-Rectangular Reward-Robust MDPs via Frequency Regularization
by: Gadot, Uri, et al.
Published: (2023) -
Simulus: Combining Improvements in Sample-Efficient World Model Agents
by: Cohen, Lior, et al.
Published: (2025)