O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yao, Yi, Zhu, He, Wang, Piaohong, Ren, Jincheng, Yang, Xinlong, Chen, Qianben, Li, Xiaowan, Shi, Dingfeng, Li, Jiaxian, Wang, Qiexiang, Wang, Sinuo, Liu, Xinpeng, Wu, Jiaqi, Liu, Minghao, Zhou, Wangchunshu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912806730727424
author Yao, Yi
Zhu, He
Wang, Piaohong
Ren, Jincheng
Yang, Xinlong
Chen, Qianben
Li, Xiaowan
Shi, Dingfeng
Li, Jiaxian
Wang, Qiexiang
Wang, Sinuo
Liu, Xinpeng
Wu, Jiaqi
Liu, Minghao
Zhou, Wangchunshu
author_facet Yao, Yi
Zhu, He
Wang, Piaohong
Ren, Jincheng
Yang, Xinlong
Chen, Qianben
Li, Xiaowan
Shi, Dingfeng
Li, Jiaxian
Wang, Qiexiang
Wang, Sinuo
Liu, Xinpeng
Wu, Jiaqi
Liu, Minghao
Zhou, Wangchunshu
contents The performance gap between closed-source and open-source large language models (LLMs) is largely attributed to disparities in access to high-quality training data. To bridge this gap, we introduce a novel framework for the automated synthesis of sophisticated, research-grade instructional data. Our approach centers on a multi-agent workflow where collaborative AI agents simulate complex tool-integrated reasoning to generate diverse and high-fidelity data end-to-end. Leveraging this synthesized data, we develop a two-stage training strategy that integrates supervised fine-tuning with a novel reinforcement learning method, designed to maximize model alignment and capability. Extensive experiments demonstrate that our framework empowers open-source models across multiple scales, enabling them to achieve new state-of-the-art performance on the major deep research benchmark. This work provides a scalable and effective pathway for advancing open-source LLMs without relying on proprietary data or models.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03743
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL
Yao, Yi
Zhu, He
Wang, Piaohong
Ren, Jincheng
Yang, Xinlong
Chen, Qianben
Li, Xiaowan
Shi, Dingfeng
Li, Jiaxian
Wang, Qiexiang
Wang, Sinuo
Liu, Xinpeng
Wu, Jiaqi
Liu, Minghao
Zhou, Wangchunshu
Computation and Language
Artificial Intelligence
The performance gap between closed-source and open-source large language models (LLMs) is largely attributed to disparities in access to high-quality training data. To bridge this gap, we introduce a novel framework for the automated synthesis of sophisticated, research-grade instructional data. Our approach centers on a multi-agent workflow where collaborative AI agents simulate complex tool-integrated reasoning to generate diverse and high-fidelity data end-to-end. Leveraging this synthesized data, we develop a two-stage training strategy that integrates supervised fine-tuning with a novel reinforcement learning method, designed to maximize model alignment and capability. Extensive experiments demonstrate that our framework empowers open-source models across multiple scales, enabling them to achieve new state-of-the-art performance on the major deep research benchmark. This work provides a scalable and effective pathway for advancing open-source LLMs without relying on proprietary data or models.
title O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.03743