WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Junteng, Li, Yunji, Zhang, Chi, Li, Jingyang, Chen, Aili, Ji, Ke, Cheng, Weiyu, Wu, Zijia, Du, Chengyu, Xu, Qidi, Song, Jiayuan, Zhu, Zhengmao, Chen, Wenhu, Zhao, Pengyu, He, Junxian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918148604690432
author Liu, Junteng
Li, Yunji
Zhang, Chi
Li, Jingyang
Chen, Aili
Ji, Ke
Cheng, Weiyu
Wu, Zijia
Du, Chengyu
Xu, Qidi
Song, Jiayuan
Zhu, Zhengmao
Chen, Wenhu
Zhao, Pengyu
He, Junxian
author_facet Liu, Junteng
Li, Yunji
Zhang, Chi
Li, Jingyang
Chen, Aili
Ji, Ke
Cheng, Weiyu
Wu, Zijia
Du, Chengyu
Xu, Qidi
Song, Jiayuan
Zhu, Zhengmao
Chen, Wenhu
Zhao, Pengyu
He, Junxian
contents The paradigm of Large Language Models (LLMs) has increasingly shifted toward agentic applications, where web browsing capabilities are fundamental for retrieving information from diverse online sources. However, existing open-source web agents either demonstrate limited information-seeking abilities on complex tasks or lack transparent implementations. In this work, we identify that the key challenge lies in the scarcity of challenging data for information seeking. To address this limitation, we introduce WebExplorer: a systematic data generation approach using model-based exploration and iterative, long-to-short query evolution. This method creates challenging query-answer pairs that require multi-step reasoning and complex web navigation. By leveraging our curated high-quality dataset, we successfully develop advanced web agent WebExplorer-8B through supervised fine-tuning followed by reinforcement learning. Our model supports 128K context length and up to 100 tool calling turns, enabling long-horizon problem solving. Across diverse information-seeking benchmarks, WebExplorer-8B achieves the state-of-the-art performance at its scale. Notably, as an 8B-sized model, WebExplorer-8B is able to effectively search over an average of 16 turns after RL training, achieving higher accuracy than WebSailor-72B on BrowseComp-en/zh and attaining the best performance among models up to 100B parameters on WebWalkerQA and FRAMES. Beyond these information-seeking tasks, our model also achieves strong generalization on the HLE benchmark even though it is only trained on knowledge-intensive QA data. These results highlight our approach as a practical path toward long-horizon web agents.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06501
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
Liu, Junteng
Li, Yunji
Zhang, Chi
Li, Jingyang
Chen, Aili
Ji, Ke
Cheng, Weiyu
Wu, Zijia
Du, Chengyu
Xu, Qidi
Song, Jiayuan
Zhu, Zhengmao
Chen, Wenhu
Zhao, Pengyu
He, Junxian
Computation and Language
The paradigm of Large Language Models (LLMs) has increasingly shifted toward agentic applications, where web browsing capabilities are fundamental for retrieving information from diverse online sources. However, existing open-source web agents either demonstrate limited information-seeking abilities on complex tasks or lack transparent implementations. In this work, we identify that the key challenge lies in the scarcity of challenging data for information seeking. To address this limitation, we introduce WebExplorer: a systematic data generation approach using model-based exploration and iterative, long-to-short query evolution. This method creates challenging query-answer pairs that require multi-step reasoning and complex web navigation. By leveraging our curated high-quality dataset, we successfully develop advanced web agent WebExplorer-8B through supervised fine-tuning followed by reinforcement learning. Our model supports 128K context length and up to 100 tool calling turns, enabling long-horizon problem solving. Across diverse information-seeking benchmarks, WebExplorer-8B achieves the state-of-the-art performance at its scale. Notably, as an 8B-sized model, WebExplorer-8B is able to effectively search over an average of 16 turns after RL training, achieving higher accuracy than WebSailor-72B on BrowseComp-en/zh and attaining the best performance among models up to 100B parameters on WebWalkerQA and FRAMES. Beyond these information-seeking tasks, our model also achieves strong generalization on the HLE benchmark even though it is only trained on knowledge-intensive QA data. These results highlight our approach as a practical path toward long-horizon web agents.
title WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
topic Computation and Language
url https://arxiv.org/abs/2509.06501