WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Guanzhong, Yang, Zhen, Liu, Jinxin, Xu, Bin, Hou, Lei, Li, Juanzi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918165090402304
author He, Guanzhong
Yang, Zhen
Liu, Jinxin
Xu, Bin
Hou, Lei
Li, Juanzi
author_facet He, Guanzhong
Yang, Zhen
Liu, Jinxin
Xu, Bin
Hou, Lei
Li, Juanzi
contents Search agents have achieved significant advancements in enabling intelligent information retrieval and decision-making within interactive environments. Although reinforcement learning has been employed to train agentic models capable of more dynamic interactive retrieval, existing methods are limited by shallow tool-use depth and the accumulation of errors over multiple iterative interactions. In this paper, we present WebSeer, a more intelligent search agent trained via reinforcement learning enhanced with a self-reflection mechanism. Specifically, we construct a large dataset annotated with reflection patterns and design a two-stage training framework that unifies cold start and reinforcement learning within the self-reflection paradigm for real-world web-based environments, which enables the model to generate longer and more reflective tool-use trajectories. Our approach substantially extends tool-use chains and improves answer accuracy. Using a single 14B model, we achieve state-of-the-art results on HotpotQA and SimpleQA, with accuracies of 72.3% and 90.0%, respectively, and demonstrate strong generalization to out-of-distribution datasets. The code is available at https://github.com/99hgz/WebSeer
format Preprint
id arxiv_https___arxiv_org_abs_2510_18798
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
He, Guanzhong
Yang, Zhen
Liu, Jinxin
Xu, Bin
Hou, Lei
Li, Juanzi
Computation and Language
Search agents have achieved significant advancements in enabling intelligent information retrieval and decision-making within interactive environments. Although reinforcement learning has been employed to train agentic models capable of more dynamic interactive retrieval, existing methods are limited by shallow tool-use depth and the accumulation of errors over multiple iterative interactions. In this paper, we present WebSeer, a more intelligent search agent trained via reinforcement learning enhanced with a self-reflection mechanism. Specifically, we construct a large dataset annotated with reflection patterns and design a two-stage training framework that unifies cold start and reinforcement learning within the self-reflection paradigm for real-world web-based environments, which enables the model to generate longer and more reflective tool-use trajectories. Our approach substantially extends tool-use chains and improves answer accuracy. Using a single 14B model, we achieve state-of-the-art results on HotpotQA and SimpleQA, with accuracies of 72.3% and 90.0%, respectively, and demonstrate strong generalization to out-of-distribution datasets. The code is available at https://github.com/99hgz/WebSeer
title WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
topic Computation and Language
url https://arxiv.org/abs/2510.18798