Plan Before Search: Search Agents Need Plan

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Qian, Zhipeng, Liang, Zihan, Ma, Yufei, Chen, Ben, Dai, Huangyu, Ji, Jiayi, Lei, Chenyi, Ou, Wenwu, Sun, Xiaoshuai, Hou, Qibin
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916055591419904
author Qian, Zhipeng
Liang, Zihan
Ma, Yufei
Chen, Ben
Dai, Huangyu
Ji, Jiayi
Lei, Chenyi
Ou, Wenwu
Sun, Xiaoshuai
Hou, Qibin
author_facet Qian, Zhipeng
Liang, Zihan
Ma, Yufei
Chen, Ben
Dai, Huangyu
Ji, Jiayi
Lei, Chenyi
Ou, Wenwu
Sun, Xiaoshuai
Hou, Qibin
contents Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a stronger model. However, this paradigm overlooks two fundamental factors: the dependency structure among sub-skills, and the possibility that distillation is not the only route to capability acquisition. We study this through Plan, a structured agentic behavior for multi-hop retrieval that decomposes a question into ordered sub-questions before any retrieval is performed, so that each search step can be anchored to a pre-designed sub-question instead of drifting under the influence of partially relevant documents retrieved earlier. However, across three model families spanning 3B to 14B parameters, we find that an identical reward signal induces qualitatively different RL failure modes. This phenomenon indicates that successful training hinges not only on reward design but also on model-specific feasibility conditions: sufficient initial entropy, training stability, and prerequisite sub-skills. Motivated by this, we propose a self-bootstrapping paradigm in which a small-scale seed model generates filtered trajectories that activate Plan in any target model, eliminating the need for distillation from an external stronger model. Our pipeline activates Plan across every tested model and consistently outperforms competitive baselines on multi-hop QA benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_28354
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Plan Before Search: Search Agents Need Plan
Qian, Zhipeng
Liang, Zihan
Ma, Yufei
Chen, Ben
Dai, Huangyu
Ji, Jiayi
Lei, Chenyi
Ou, Wenwu
Sun, Xiaoshuai
Hou, Qibin
Artificial Intelligence
Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a stronger model. However, this paradigm overlooks two fundamental factors: the dependency structure among sub-skills, and the possibility that distillation is not the only route to capability acquisition. We study this through Plan, a structured agentic behavior for multi-hop retrieval that decomposes a question into ordered sub-questions before any retrieval is performed, so that each search step can be anchored to a pre-designed sub-question instead of drifting under the influence of partially relevant documents retrieved earlier. However, across three model families spanning 3B to 14B parameters, we find that an identical reward signal induces qualitatively different RL failure modes. This phenomenon indicates that successful training hinges not only on reward design but also on model-specific feasibility conditions: sufficient initial entropy, training stability, and prerequisite sub-skills. Motivated by this, we propose a self-bootstrapping paradigm in which a small-scale seed model generates filtered trajectories that activate Plan in any target model, eliminating the need for distillation from an external stronger model. Our pipeline activates Plan across every tested model and consistently outperforms competitive baselines on multi-hop QA benchmarks.
title Plan Before Search: Search Agents Need Plan
topic Artificial Intelligence
url https://arxiv.org/abs/2605.28354