OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908845674070016 |
|---|---|
| author | Zhou, Yuhang Zheng, Kai Chen, Qiguang Hu, Mengkang Sun, Qingfeng Xu, Can Chen, Jingjing |
| author_facet | Zhou, Yuhang Zheng, Kai Chen, Qiguang Hu, Mengkang Sun, Qingfeng Xu, Can Chen, Jingjing |
| contents | Deep research agents have shown remarkable potential in handling long-horizon tasks. However, state-of-the-art performance typically relies on online reinforcement learning (RL), which is financially expensive due to extensive API calls. While offline training offers a more efficient alternative, its progress is hindered by the scarcity of high-quality research trajectories. In this paper, we demonstrate that expensive online reinforcement learning is not all you need to build powerful research agents. To bridge this gap, we introduce a fully open-source suite designed for effective offline training. Our core contributions include DeepForge, a ready-to-use task synthesis framework that generates large-scale research queries without heavy preprocessing; and a curated collection of 66k QA pairs, 33k SFT trajectories, and 21k DPO pairs. Leveraging these resources, we train OffSeeker (8B), a model developed entirely offline. Extensive evaluations across six benchmarks show that OffSeeker not only leads among similar-sized agents but also remains competitive with 30B-parameter systems trained via heavy online RL. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_18467 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents Zhou, Yuhang Zheng, Kai Chen, Qiguang Hu, Mengkang Sun, Qingfeng Xu, Can Chen, Jingjing Artificial Intelligence Machine Learning Deep research agents have shown remarkable potential in handling long-horizon tasks. However, state-of-the-art performance typically relies on online reinforcement learning (RL), which is financially expensive due to extensive API calls. While offline training offers a more efficient alternative, its progress is hindered by the scarcity of high-quality research trajectories. In this paper, we demonstrate that expensive online reinforcement learning is not all you need to build powerful research agents. To bridge this gap, we introduce a fully open-source suite designed for effective offline training. Our core contributions include DeepForge, a ready-to-use task synthesis framework that generates large-scale research queries without heavy preprocessing; and a curated collection of 66k QA pairs, 33k SFT trajectories, and 21k DPO pairs. Leveraging these resources, we train OffSeeker (8B), a model developed entirely offline. Extensive evaluations across six benchmarks show that OffSeeker not only leads among similar-sized agents but also remains competitive with 30B-parameter systems trained via heavy online RL. |
| title | OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents |
| topic | Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2601.18467 |