How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yinuo, Lu, Shuo, Cheng, Jianjie, Wang, Meng, Xie, Qianlong, Wang, Xingxing, He, Ran, Liang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
by: Lu, Shuo, et al.
Published: (2025)
by: Lu, Shuo, et al.
Published: (2025)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
by: Yu, Yongcan, et al.
Published: (2026)
by: Yu, Yongcan, et al.
Published: (2026)
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
by: Lu, Shuo, et al.
Published: (2026)
by: Lu, Shuo, et al.
Published: (2026)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
by: Sheng, Lijun, et al.
Published: (2025)
by: Sheng, Lijun, et al.
Published: (2025)
Off-Policy Primal-Dual Safe Reinforcement Learning
by: Wu, Zifan, et al.
Published: (2024)
by: Wu, Zifan, et al.
Published: (2024)
RIA: A Ranking-Infused Approach for Optimized listwise CTR Prediction
by: Zhang, Guoxiao, et al.
Published: (2025)
by: Zhang, Guoxiao, et al.
Published: (2025)
FITRep: Attention-Guided Item Representation via MLLMs
by: Zhang, Guoxiao, et al.
Published: (2025)
by: Zhang, Guoxiao, et al.
Published: (2025)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
by: Ding, Ruiyi, et al.
Published: (2026)
by: Ding, Ruiyi, et al.
Published: (2026)
Generative Large-Scale Pre-trained Models for Automated Ad Bidding Optimization
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense
by: Lan, Qianlong, et al.
Published: (2026)
by: Lan, Qianlong, et al.
Published: (2026)
Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?
by: Zhao, Yibo, et al.
Published: (2026)
by: Zhao, Yibo, et al.
Published: (2026)
Beyond Single Slot: Joint Optimization for Multi-Slot Guaranteed Display Advertising
by: Zhang, Zhaoqi, et al.
Published: (2026)
by: Zhang, Zhaoqi, et al.
Published: (2026)
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards
by: Yan, Wentao, et al.
Published: (2026)
by: Yan, Wentao, et al.
Published: (2026)
Comment on ‘Aquablation for benign prostatic hyperplasia: real‐world prostate size relevance and bleeding events across 6 years’
by: Shuo Lin, et al.
Published: (2026)
by: Shuo Lin, et al.
Published: (2026)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
by: Sun, Wei, et al.
Published: (2025)
by: Sun, Wei, et al.
Published: (2025)
Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Training
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
by: Wang, Yiding, et al.
Published: (2025)
by: Wang, Yiding, et al.
Published: (2025)
PaperScout: An Autonomous Agent for Academic Paper Search with Process-Aware Sequence-Level Policy Optimization
by: Pan, Tingyue, et al.
Published: (2026)
by: Pan, Tingyue, et al.
Published: (2026)
Density properties of orbits for a hypercyclic operator on a Banach space
by: Li, Jian, et al.
Published: (2025)
by: Li, Jian, et al.
Published: (2025)
Mean Li-Yorke chaos for a sequence of operators on Banach spaces
by: Li, Jian, et al.
Published: (2025)
by: Li, Jian, et al.
Published: (2025)
WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents
by: Wang, Xilong, et al.
Published: (2026)
by: Wang, Xilong, et al.
Published: (2026)
Sample Correlation for Fingerprinting Deep Face Recognition
by: Guan, Jiyang, et al.
Published: (2024)
by: Guan, Jiyang, et al.
Published: (2024)
Silent Egress: When Implicit Prompt Injection Makes LLM Agents Leak Without a Trace
by: Lan, Qianlong, et al.
Published: (2026)
by: Lan, Qianlong, et al.
Published: (2026)
How to Train Your Metamorphic Deep Neural Network
by: Sommariva, Thomas, et al.
Published: (2025)
by: Sommariva, Thomas, et al.
Published: (2025)
SE-Search: Self-Evolving Search Agent via Memory and Dense Reward
by: Li, Jian, et al.
Published: (2026)
by: Li, Jian, et al.
Published: (2026)
Bootstrap Your Own Context Length
by: Wang, Liang, et al.
Published: (2024)
by: Wang, Liang, et al.
Published: (2024)
HiBid: A Cross-Channel Constrained Bidding System with Budget Allocation by Hierarchical Offline Deep Reinforcement Learning
by: Wang, Hao, et al.
Published: (2023)
by: Wang, Hao, et al.
Published: (2023)
HiconAgent: History Context-aware Policy Optimization for GUI Agents
by: Zhou, Xurui, et al.
Published: (2025)
by: Zhou, Xurui, et al.
Published: (2025)
WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents
by: Liu, Yinuo, et al.
Published: (2025)
by: Liu, Yinuo, et al.
Published: (2025)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
by: Hao, Ruijie, et al.
Published: (2026)
by: Hao, Ruijie, et al.
Published: (2026)
Frustratingly Easy Feature Reconstruction for Out-of-Distribution Detection
by: Wang, Yingsheng, et al.
Published: (2025)
by: Wang, Yingsheng, et al.
Published: (2025)
Cycle-Consistent Search: Question Reconstructability as a Proxy Reward for Search Agent Training
by: An, Sohyun, et al.
Published: (2026)
by: An, Sohyun, et al.
Published: (2026)
QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
by: Xie, Jian, et al.
Published: (2026)
by: Xie, Jian, et al.
Published: (2026)
RL-MPCA: A Reinforcement Learning Based Multi-Phase Computation Allocation Approach for Recommender Systems
by: Zhou, Jiahong, et al.
Published: (2023)
by: Zhou, Jiahong, et al.
Published: (2023)
Safe Offline Reinforcement Learning with Real-Time Budget Constraints
by: Lin, Qian, et al.
Published: (2023)
by: Lin, Qian, et al.
Published: (2023)
MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt Optimization
by: Zhang, Jian, et al.
Published: (2025)
by: Zhang, Jian, et al.
Published: (2025)
Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Comparative Statistical Analysis of Prompt and Afterglow X-Ray Flares in Gamma-Ray Bursts: Insights into Extended Central Engine Activity
by: Ma, Yinuo, et al.
Published: (2025)
by: Ma, Yinuo, et al.
Published: (2025)
Similar Items
-
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
by: Lu, Shuo, et al.
Published: (2025) -
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
by: Yu, Yongcan, et al.
Published: (2025) -
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning
by: Yu, Yongcan, et al.
Published: (2026) -
Do MLLMs Really Understand Space? A Mathematical Reasoning Evaluation
by: Lu, Shuo, et al.
Published: (2026) -
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)