Generalized Nested Rollout Policy Adaptation with Limited Repetitions
Fuente:
arXiv
Saved in:
| Main Author: | Cazenave, Tristan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Bias Generalized Rollout Policy Adaptation on the Flexible Job-Shop Scheduling Problem
by: Kobrosly, Lotfi, et al.
Published: (2025)
by: Kobrosly, Lotfi, et al.
Published: (2025)
Monte Carlo Search Algorithms Discovering Monte Carlo Tree Search Exploration Terms
by: Cazenave, Tristan
Published: (2024)
by: Cazenave, Tristan
Published: (2024)
Learning a Prior for Monte Carlo Search by Replaying Solutions to Combinatorial Problems
by: Cazenave, Tristan
Published: (2024)
by: Cazenave, Tristan
Published: (2024)
Eterna is Solved
by: Cazenave, Tristan
Published: (2025)
by: Cazenave, Tristan
Published: (2025)
Monte Carlo Permutation Search
by: Cazenave, Tristan
Published: (2025)
by: Cazenave, Tristan
Published: (2025)
Generalized Rapid Action Value Estimation in Memory-Constrained Environments
by: Rautureau, Aloïs, et al.
Published: (2026)
by: Rautureau, Aloïs, et al.
Published: (2026)
A General Highly Accurate Online Planning Method Integrating Large Language Models into Nested Rollout Policy Adaptation for Dialogue Tasks
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Minibal: Balanced Game-Playing Without Opponent Modeling
by: Cohen-Solal, Quentin, et al.
Published: (2026)
by: Cohen-Solal, Quentin, et al.
Published: (2026)
Minimax Strikes Back
by: Cohen-Solal, Quentin, et al.
Published: (2020)
by: Cohen-Solal, Quentin, et al.
Published: (2020)
On some improvements to Unbounded Minimax
by: Cohen-Solal, Quentin, et al.
Published: (2025)
by: Cohen-Solal, Quentin, et al.
Published: (2025)
ACCORD: Autoregressive Constraint-satisfying Generation for COmbinatorial Optimization with Routing and Dynamic attention
by: Abgaryan, Henrik, et al.
Published: (2025)
by: Abgaryan, Henrik, et al.
Published: (2025)
Perfect Information Monte Carlo with Postponing Reasoning
by: Arjonilla, Jérôme, et al.
Published: (2024)
by: Arjonilla, Jérôme, et al.
Published: (2024)
Enhancing Reinforcement Learning Through Guided Search
by: Arjonilla, Jérôme, et al.
Published: (2024)
by: Arjonilla, Jérôme, et al.
Published: (2024)
LLMs can Schedule
by: Abgaryan, Henrik, et al.
Published: (2024)
by: Abgaryan, Henrik, et al.
Published: (2024)
Monte Carlo Graph Coloring
by: Cazenave, Tristan, et al.
Published: (2025)
by: Cazenave, Tristan, et al.
Published: (2025)
SpinGPT: A Large-Language-Model Approach to Playing Poker Correctly
by: Maugin, Narada, et al.
Published: (2025)
by: Maugin, Narada, et al.
Published: (2025)
BeeRNA: tertiary structure-based RNA inverse folding using Artificial Bee Colony
by: Mlaweh, Mehyar, et al.
Published: (2025)
by: Mlaweh, Mehyar, et al.
Published: (2025)
Starjob: Dataset for LLM-Driven Job Shop Scheduling
by: Abgaryan, Henrik, et al.
Published: (2025)
by: Abgaryan, Henrik, et al.
Published: (2025)
Pareto-NRPA: A Novel Monte-Carlo Search Algorithm for Multi-Objective Optimization
by: Lallouet, Noé, et al.
Published: (2025)
by: Lallouet, Noé, et al.
Published: (2025)
Mixture of Public and Private Distributions in Imperfect Information Games
by: Arjonilla, Jérôme, et al.
Published: (2024)
by: Arjonilla, Jérôme, et al.
Published: (2024)
Deep Reinforcement Learning for 5*5 Multiplayer Go
by: Driss, Brahim, et al.
Published: (2024)
by: Driss, Brahim, et al.
Published: (2024)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
by: Heuillet, Maxime, et al.
Published: (2025)
by: Heuillet, Maxime, et al.
Published: (2025)
Refutation of Spectral Graph Theory Conjectures with Search Algorithms)
by: Roucairol, Milo, et al.
Published: (2024)
by: Roucairol, Milo, et al.
Published: (2024)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
by: Liu, Bingshuai, et al.
Published: (2025)
by: Liu, Bingshuai, et al.
Published: (2025)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
by: Yoo, Seungwoo, et al.
Published: (2026)
by: Yoo, Seungwoo, et al.
Published: (2026)
How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization
by: Fang, Yangyi, et al.
Published: (2026)
by: Fang, Yangyi, et al.
Published: (2026)
Portfolio Reinforcement Learning with Scenario-Context Rollout
by: Bendatu, Vanya Priscillia, et al.
Published: (2026)
by: Bendatu, Vanya Priscillia, et al.
Published: (2026)
Agentic Flow Steering and Parallel Rollout Search for Spatially Grounded Text-to-Image Generation
by: Chen, Ping, et al.
Published: (2026)
by: Chen, Ping, et al.
Published: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
by: Kim, Youngeun
Published: (2026)
by: Kim, Youngeun
Published: (2026)
Rollout Cards: A Reproducibility Standard for Agent Research
by: Masters, Charlie, et al.
Published: (2026)
by: Masters, Charlie, et al.
Published: (2026)
Maximum Entropy Exploration Without the Rollouts
by: Adamczyk, Jacob, et al.
Published: (2026)
by: Adamczyk, Jacob, et al.
Published: (2026)
APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
by: Zhou, Yuzhen, et al.
Published: (2025)
by: Zhou, Yuzhen, et al.
Published: (2025)
AnchorDrive: LLM Scenario Rollout with Anchor-Guided Diffusion Regeneration for Safety-Critical Scenario Generation
by: Jiang, Zhulin, et al.
Published: (2026)
by: Jiang, Zhulin, et al.
Published: (2026)
Nested Graph Pseudo-Label Refinement for Noisy Label Domain Adaptation Learning
by: Wang, Yingxu, et al.
Published: (2025)
by: Wang, Yingxu, et al.
Published: (2025)
CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback
by: Chen, Bin, et al.
Published: (2026)
by: Chen, Bin, et al.
Published: (2026)
Controlling Repetition in Protein Language Models
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
Continual Learning in the Presence of Repetition
by: Hemati, Hamed, et al.
Published: (2024)
by: Hemati, Hamed, et al.
Published: (2024)
Federated Nested Learning: Collaborative Training of Self-Referential Memories for Test-Time Adaptation
by: Chen, Hong, et al.
Published: (2026)
by: Chen, Hong, et al.
Published: (2026)
Code Copycat Conundrum: Demystifying Repetition in LLM-based Code Generation
by: Liu, Mingwei, et al.
Published: (2025)
by: Liu, Mingwei, et al.
Published: (2025)
Similar Items
-
Adaptive Bias Generalized Rollout Policy Adaptation on the Flexible Job-Shop Scheduling Problem
by: Kobrosly, Lotfi, et al.
Published: (2025) -
Monte Carlo Search Algorithms Discovering Monte Carlo Tree Search Exploration Terms
by: Cazenave, Tristan
Published: (2024) -
Learning a Prior for Monte Carlo Search by Replaying Solutions to Combinatorial Problems
by: Cazenave, Tristan
Published: (2024) -
Eterna is Solved
by: Cazenave, Tristan
Published: (2025) -
Monte Carlo Permutation Search
by: Cazenave, Tristan
Published: (2025)