Reinforcement Learning Foundations for Deep Research Systems: A Survey

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenjun, Chen, Zhi, Lin, Jingru, Cao, Hannan, Han, Wei, Liang, Sheng, Zhang, Zhi, Dong, Kuicai, Li, Dexun, Zhang, Chen, Liu, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917062186631168
author Li, Wenjun
Chen, Zhi
Lin, Jingru
Cao, Hannan
Han, Wei
Liang, Sheng
Zhang, Zhi
Dong, Kuicai
Li, Dexun
Zhang, Chen
Liu, Yong
author_facet Li, Wenjun
Chen, Zhi
Lin, Jingru
Cao, Hannan
Han, Wei
Liang, Sheng
Zhang, Zhi
Dong, Kuicai
Li, Dexun
Zhang, Chen
Liu, Yong
contents Deep research systems, agentic AI that solve complex, multi-step tasks by coordinating reasoning, search across the open web and user files, and tool use, are moving toward hierarchical deployments with a Planner, Coordinator, and Executors. In practice, training entire stacks end-to-end remains impractical, so most work trains a single planner connected to core tools such as search, browsing, and code. While SFT imparts protocol fidelity, it suffers from imitation and exposure biases and underuses environment feedback. Preference alignment methods such as DPO are schema and proxy-dependent, off-policy, and weak for long-horizon credit assignment and multi-objective trade-offs. A further limitation of SFT and DPO is their reliance on human defined decision points and subskills through schema design and labeled comparisons. Reinforcement learning aligns with closed-loop, tool-interaction research by optimizing trajectory-level policies, enabling exploration, recovery behaviors, and principled credit assignment, and it reduces dependence on such human priors and rater biases. This survey is, to our knowledge, the first dedicated to the RL foundations of deep research systems. It systematizes recent work along three axes: (i) data synthesis and curation; (ii) RL methods for agentic research covering stability, sample efficiency, long context handling, reward and credit design, multi-objective optimization, and multimodal integration; and (iii) agentic RL training systems and frameworks. We also cover agent architecture and coordination, as well as evaluation and benchmarks, including recent QA, VQA, long-form synthesis, and domain-grounded, tool-interaction tasks. We distill recurring patterns, surface infrastructure bottlenecks, and offer practical guidance for training robust, transparent deep research agents with RL.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06733
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning Foundations for Deep Research Systems: A Survey
Li, Wenjun
Chen, Zhi
Lin, Jingru
Cao, Hannan
Han, Wei
Liang, Sheng
Zhang, Zhi
Dong, Kuicai
Li, Dexun
Zhang, Chen
Liu, Yong
Artificial Intelligence
Computation and Language
Information Retrieval
Deep research systems, agentic AI that solve complex, multi-step tasks by coordinating reasoning, search across the open web and user files, and tool use, are moving toward hierarchical deployments with a Planner, Coordinator, and Executors. In practice, training entire stacks end-to-end remains impractical, so most work trains a single planner connected to core tools such as search, browsing, and code. While SFT imparts protocol fidelity, it suffers from imitation and exposure biases and underuses environment feedback. Preference alignment methods such as DPO are schema and proxy-dependent, off-policy, and weak for long-horizon credit assignment and multi-objective trade-offs. A further limitation of SFT and DPO is their reliance on human defined decision points and subskills through schema design and labeled comparisons. Reinforcement learning aligns with closed-loop, tool-interaction research by optimizing trajectory-level policies, enabling exploration, recovery behaviors, and principled credit assignment, and it reduces dependence on such human priors and rater biases. This survey is, to our knowledge, the first dedicated to the RL foundations of deep research systems. It systematizes recent work along three axes: (i) data synthesis and curation; (ii) RL methods for agentic research covering stability, sample efficiency, long context handling, reward and credit design, multi-objective optimization, and multimodal integration; and (iii) agentic RL training systems and frameworks. We also cover agent architecture and coordination, as well as evaluation and benchmarks, including recent QA, VQA, long-form synthesis, and domain-grounded, tool-interaction tasks. We distill recurring patterns, surface infrastructure bottlenecks, and offer practical guidance for training robust, transparent deep research agents with RL.
title Reinforcement Learning Foundations for Deep Research Systems: A Survey
topic Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2509.06733