Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liang, Zihan, Ma, Yufei, Chen, Ben, Qian, Zhipeng, Zhang, Xuxin, Dai, Huangyu, Mao, Lingtao
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918524555886592
author Liang, Zihan
Ma, Yufei
Chen, Ben
Qian, Zhipeng
Zhang, Xuxin
Dai, Huangyu
Mao, Lingtao
author_facet Liang, Zihan
Ma, Yufei
Chen, Ben
Qian, Zhipeng
Zhang, Xuxin
Dai, Huangyu
Mao, Lingtao
contents Post-training has become the dominant recipe for turning a language model into a competent search-augmented reasoning agent. A line of recent work pushes its performance further by adding elaborate machinery on top of this standard pipeline. These augmentations import external supervision from stronger external systems, attach auxiliary modules such as process reward models or retrospective critics, restructure the rollout itself with tree search or multi-stage curricula, or shape the reward with hand-crafted bonuses and penalties. Each addition delivers a measurable gain, but each also inflates the training pipeline and ties the recipe to resources or designs that may not always be available. We take a step back and ask whether any of this machinery is actually necessary, and propose Search-E1, a self-evolution method that lets a search-augmented agent improve through only vanilla GRPO interleaved with on-policy self-distillation (OPSD). After each GRPO round, the policy rolls out on its own training questions. A token-level forward KL objective then aligns the policy's inference-time distribution to its own distribution under a privileged context that exposes a more efficient sibling trajectory. Despite this simplicity, the procedure naturally provides dense per-step supervision. On seven QA benchmarks, Search-E1 reaches 0.440 average EM with Qwen2.5-3B, surpassing all open-source baselines at both scales. Code and complete version will be made public soon.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22511
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning
Liang, Zihan
Ma, Yufei
Chen, Ben
Qian, Zhipeng
Zhang, Xuxin
Dai, Huangyu
Mao, Lingtao
Artificial Intelligence
Computation and Language
Information Retrieval
Post-training has become the dominant recipe for turning a language model into a competent search-augmented reasoning agent. A line of recent work pushes its performance further by adding elaborate machinery on top of this standard pipeline. These augmentations import external supervision from stronger external systems, attach auxiliary modules such as process reward models or retrospective critics, restructure the rollout itself with tree search or multi-stage curricula, or shape the reward with hand-crafted bonuses and penalties. Each addition delivers a measurable gain, but each also inflates the training pipeline and ties the recipe to resources or designs that may not always be available. We take a step back and ask whether any of this machinery is actually necessary, and propose Search-E1, a self-evolution method that lets a search-augmented agent improve through only vanilla GRPO interleaved with on-policy self-distillation (OPSD). After each GRPO round, the policy rolls out on its own training questions. A token-level forward KL objective then aligns the policy's inference-time distribution to its own distribution under a privileged context that exposes a more efficient sibling trajectory. Despite this simplicity, the procedure naturally provides dense per-step supervision. On seven QA benchmarks, Search-E1 reaches 0.440 average EM with Qwen2.5-3B, surpassing all open-source baselines at both scales. Code and complete version will be made public soon.
title Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoning
topic Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2605.22511