Planner-Centric Reinforcement Learning for Deep Research with Structure-Aware Reward
Fuente:
arXiv
Guardado en:
| Autores principales: | Hussain, Mustafa Anis, Wu, Xinle, Lu, Yao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Autonomous Memory Agents
por: Wu, Xinle, et al.
Publicado: (2026)
por: Wu, Xinle, et al.
Publicado: (2026)
Reward Model Routing in Alignment
por: Wu, Xinle, et al.
Publicado: (2025)
por: Wu, Xinle, et al.
Publicado: (2025)
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
por: Zhang, Rui, et al.
Publicado: (2026)
por: Zhang, Rui, et al.
Publicado: (2026)
Automatic Configuration of LLM Post-Training Pipelines
por: Chwa, Channe, et al.
Publicado: (2026)
por: Chwa, Channe, et al.
Publicado: (2026)
Transformable Gaussian Reward Function for Socially-Aware Navigation with Deep Reinforcement Learning
por: Kim, Jinyeob, et al.
Publicado: (2024)
por: Kim, Jinyeob, et al.
Publicado: (2024)
CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving
por: Guo, Yihong, et al.
Publicado: (2026)
por: Guo, Yihong, et al.
Publicado: (2026)
DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
por: Fan, Wei, et al.
Publicado: (2025)
por: Fan, Wei, et al.
Publicado: (2025)
Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
por: Duan, Xintong, et al.
Publicado: (2025)
por: Duan, Xintong, et al.
Publicado: (2025)
Deep Reinforcement Learning via Object-Centric Attention
por: Blüml, Jannis, et al.
Publicado: (2025)
por: Blüml, Jannis, et al.
Publicado: (2025)
Object-Centric World Models for Causality-Aware Reinforcement Learning
por: Nishimoto, Yosuke, et al.
Publicado: (2025)
por: Nishimoto, Yosuke, et al.
Publicado: (2025)
Reward Training Wheels: Adaptive Auxiliary Rewards for Robotics Reinforcement Learning
por: Wang, Linji, et al.
Publicado: (2025)
por: Wang, Linji, et al.
Publicado: (2025)
Reward Models in Deep Reinforcement Learning: A Survey
por: Yu, Rui, et al.
Publicado: (2025)
por: Yu, Rui, et al.
Publicado: (2025)
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search
por: Wu, Fang, et al.
Publicado: (2025)
por: Wu, Fang, et al.
Publicado: (2025)
Expert Q-learning: Deep Reinforcement Learning with Coarse State Values from Offline Expert Examples
por: Meng, Li, et al.
Publicado: (2021)
por: Meng, Li, et al.
Publicado: (2021)
Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards
por: Han, Tianyang, et al.
Publicado: (2026)
por: Han, Tianyang, et al.
Publicado: (2026)
Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards
por: Metcalf, Katherine, et al.
Publicado: (2024)
por: Metcalf, Katherine, et al.
Publicado: (2024)
Let Hybrid A* Path Planner Obey Traffic Rules: A Deep Reinforcement Learning-Based Planning Framework
por: Li, Xibo, et al.
Publicado: (2024)
por: Li, Xibo, et al.
Publicado: (2024)
Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
por: Wei, Xiaolong, et al.
Publicado: (2025)
por: Wei, Xiaolong, et al.
Publicado: (2025)
Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
por: Zhu, Siyu, et al.
Publicado: (2025)
por: Zhu, Siyu, et al.
Publicado: (2025)
Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following
por: Yang, Yuxiao, et al.
Publicado: (2024)
por: Yang, Yuxiao, et al.
Publicado: (2024)
Reward Hacking in Rubric-Based Reinforcement Learning
por: Mahmoud, Anas, et al.
Publicado: (2026)
por: Mahmoud, Anas, et al.
Publicado: (2026)
HLS-Seek: QoR-Aware Code Generation for High-Level Synthesis via Proxy Comparative Reward Reinforcement Learning
por: Zou, Qingyun, et al.
Publicado: (2026)
por: Zou, Qingyun, et al.
Publicado: (2026)
SPARK: Stepwise Process-Aware Rewards for Reference-Free Reinforcement Learning
por: Rahman, Salman, et al.
Publicado: (2025)
por: Rahman, Salman, et al.
Publicado: (2025)
PEAR: Phase Entropy Aware Reward for Efficient Reasoning
por: Huang, Chen, et al.
Publicado: (2025)
por: Huang, Chen, et al.
Publicado: (2025)
Removing Planner Bias in Goal Recognition Through Multi-Plan Dataset Generation
por: Abdelwahed, Mustafa F., et al.
Publicado: (2026)
por: Abdelwahed, Mustafa F., et al.
Publicado: (2026)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
por: Lu, Xiaodong, et al.
Publicado: (2026)
por: Lu, Xiaodong, et al.
Publicado: (2026)
Portfolio Reinforcement Learning with Scenario-Context Rollout
por: Bendatu, Vanya Priscillia, et al.
Publicado: (2026)
por: Bendatu, Vanya Priscillia, et al.
Publicado: (2026)
OccuReward: LLM-Guided Occupant-Centric Reward Shaping for Demographic Equity in Grid-Interactive Buildings
por: Zaregarizi, Shadmehr, et al.
Publicado: (2026)
por: Zaregarizi, Shadmehr, et al.
Publicado: (2026)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
por: Wang, Jiayu, et al.
Publicado: (2025)
por: Wang, Jiayu, et al.
Publicado: (2025)
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
por: Cai, Xinyan, et al.
Publicado: (2025)
por: Cai, Xinyan, et al.
Publicado: (2025)
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
por: Shen, Wei, et al.
Publicado: (2024)
por: Shen, Wei, et al.
Publicado: (2024)
Counting Reward Automata: Sample Efficient Reinforcement Learning Through the Exploitation of Reward Function Structure
por: Bester, Tristan, et al.
Publicado: (2023)
por: Bester, Tristan, et al.
Publicado: (2023)
AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
por: Mei, Lang, et al.
Publicado: (2025)
por: Mei, Lang, et al.
Publicado: (2025)
Reinforcement Learning with Robust Rubric Rewards
por: Yu, Ya-Qi, et al.
Publicado: (2026)
por: Yu, Ya-Qi, et al.
Publicado: (2026)
Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement Learning
por: Azran, Guy, et al.
Publicado: (2023)
por: Azran, Guy, et al.
Publicado: (2023)
DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
por: Zheng, Yuxiang, et al.
Publicado: (2025)
por: Zheng, Yuxiang, et al.
Publicado: (2025)
Rewarding What Matters: Step-by-Step Reinforcement Learning for Task-Oriented Dialogue
por: Du, Huifang, et al.
Publicado: (2024)
por: Du, Huifang, et al.
Publicado: (2024)
Deep Reinforcement Learning for Fault-Adaptive Routing in Eisenstein-Jacobi Interconnection Topologies
por: Charrwi, Mohammad Walid, et al.
Publicado: (2026)
por: Charrwi, Mohammad Walid, et al.
Publicado: (2026)
TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning
por: Wang, Yinuo, et al.
Publicado: (2026)
por: Wang, Yinuo, et al.
Publicado: (2026)
Reinforcement Learning with Symbolic Reward Machines
por: Krug, Thomas, et al.
Publicado: (2026)
por: Krug, Thomas, et al.
Publicado: (2026)
Ejemplares similares
-
Towards Autonomous Memory Agents
por: Wu, Xinle, et al.
Publicado: (2026) -
Reward Model Routing in Alignment
por: Wu, Xinle, et al.
Publicado: (2025) -
CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts
por: Zhang, Rui, et al.
Publicado: (2026) -
Automatic Configuration of LLM Post-Training Pipelines
por: Chwa, Channe, et al.
Publicado: (2026) -
Transformable Gaussian Reward Function for Socially-Aware Navigation with Deep Reinforcement Learning
por: Kim, Jinyeob, et al.
Publicado: (2024)