SynPlanResearch-R1: Encouraging Tool Exploration for Deep Research with Synthetic Plans

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Hansi, Li, Zoey, Gao, Yifan, Zhang, Chenwei, Pan, Xiaoman, Yang, Tao, Mo, Fengran, Lin, Jiacheng, Li, Xian, Shang, Jingbo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915845481955328
author Zeng, Hansi
Li, Zoey
Gao, Yifan
Zhang, Chenwei
Pan, Xiaoman
Yang, Tao
Mo, Fengran
Lin, Jiacheng
Li, Xian
Shang, Jingbo
author_facet Zeng, Hansi
Li, Zoey
Gao, Yifan
Zhang, Chenwei
Pan, Xiaoman
Yang, Tao
Mo, Fengran
Lin, Jiacheng
Li, Xian
Shang, Jingbo
contents Research Agents enable models to gather information from the web using tools to answer user queries, requiring them to dynamically interleave internal reasoning with tool use. While such capabilities can in principle be learned via reinforcement learning with verifiable rewards (RLVR), we observe that agents often exhibit poor exploration behaviors, including premature termination and biased tool usage. As a result, RLVR alone yields limited improvements. We propose SynPlanResearch-R1, a framework that synthesizes tool-use trajectories that encourage deeper exploration to shape exploration during cold-start supervised fine-tuning, providing a strong initialization for subsequent RL. Across seven multi-hop and open-web benchmarks, \framework improves performance by up to 6.0% on Qwen3-8B and 5.8% on Qwen3-4B backbones respectively compared to SOTA baselines. Further analyses of tool-use patterns and training dynamics compared to baselines shed light on the factors underlying these gains. Our code is publicly available at https://github.com/HansiZeng/syn-plan-research.
format Preprint
id arxiv_https___arxiv_org_abs_2603_07853
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SynPlanResearch-R1: Encouraging Tool Exploration for Deep Research with Synthetic Plans
Zeng, Hansi
Li, Zoey
Gao, Yifan
Zhang, Chenwei
Pan, Xiaoman
Yang, Tao
Mo, Fengran
Lin, Jiacheng
Li, Xian
Shang, Jingbo
Artificial Intelligence
Computation and Language
Information Retrieval
Research Agents enable models to gather information from the web using tools to answer user queries, requiring them to dynamically interleave internal reasoning with tool use. While such capabilities can in principle be learned via reinforcement learning with verifiable rewards (RLVR), we observe that agents often exhibit poor exploration behaviors, including premature termination and biased tool usage. As a result, RLVR alone yields limited improvements. We propose SynPlanResearch-R1, a framework that synthesizes tool-use trajectories that encourage deeper exploration to shape exploration during cold-start supervised fine-tuning, providing a strong initialization for subsequent RL. Across seven multi-hop and open-web benchmarks, \framework improves performance by up to 6.0% on Qwen3-8B and 5.8% on Qwen3-4B backbones respectively compared to SOTA baselines. Further analyses of tool-use patterns and training dynamics compared to baselines shed light on the factors underlying these gains. Our code is publicly available at https://github.com/HansiZeng/syn-plan-research.
title SynPlanResearch-R1: Encouraging Tool Exploration for Deep Research with Synthetic Plans
topic Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2603.07853