OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yuhang, Zheng, Kai, Chen, Qiguang, Hu, Mengkang, Sun, Qingfeng, Xu, Can, Chen, Jingjing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908845674070016
author Zhou, Yuhang
Zheng, Kai
Chen, Qiguang
Hu, Mengkang
Sun, Qingfeng
Xu, Can
Chen, Jingjing
author_facet Zhou, Yuhang
Zheng, Kai
Chen, Qiguang
Hu, Mengkang
Sun, Qingfeng
Xu, Can
Chen, Jingjing
contents Deep research agents have shown remarkable potential in handling long-horizon tasks. However, state-of-the-art performance typically relies on online reinforcement learning (RL), which is financially expensive due to extensive API calls. While offline training offers a more efficient alternative, its progress is hindered by the scarcity of high-quality research trajectories. In this paper, we demonstrate that expensive online reinforcement learning is not all you need to build powerful research agents. To bridge this gap, we introduce a fully open-source suite designed for effective offline training. Our core contributions include DeepForge, a ready-to-use task synthesis framework that generates large-scale research queries without heavy preprocessing; and a curated collection of 66k QA pairs, 33k SFT trajectories, and 21k DPO pairs. Leveraging these resources, we train OffSeeker (8B), a model developed entirely offline. Extensive evaluations across six benchmarks show that OffSeeker not only leads among similar-sized agents but also remains competitive with 30B-parameter systems trained via heavy online RL.
format Preprint
id arxiv_https___arxiv_org_abs_2601_18467
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
Zhou, Yuhang
Zheng, Kai
Chen, Qiguang
Hu, Mengkang
Sun, Qingfeng
Xu, Can
Chen, Jingjing
Artificial Intelligence
Machine Learning
Deep research agents have shown remarkable potential in handling long-horizon tasks. However, state-of-the-art performance typically relies on online reinforcement learning (RL), which is financially expensive due to extensive API calls. While offline training offers a more efficient alternative, its progress is hindered by the scarcity of high-quality research trajectories. In this paper, we demonstrate that expensive online reinforcement learning is not all you need to build powerful research agents. To bridge this gap, we introduce a fully open-source suite designed for effective offline training. Our core contributions include DeepForge, a ready-to-use task synthesis framework that generates large-scale research queries without heavy preprocessing; and a curated collection of 66k QA pairs, 33k SFT trajectories, and 21k DPO pairs. Leveraging these resources, we train OffSeeker (8B), a model developed entirely offline. Extensive evaluations across six benchmarks show that OffSeeker not only leads among similar-sized agents but also remains competitive with 30B-parameter systems trained via heavy online RL.
title OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.18467