MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866917237700427776 |
|---|---|
| author | Chen, Guoxin Qiao, Zile Wang, Wenqing Yu, Donglei Chen, Xuanzhong Sun, Hao Liao, Minpeng Fan, Kai Jiang, Yong Xie, Penguin Zhao, Wayne Xin Song, Ruihua Huang, Fei |
| author_facet | Chen, Guoxin Qiao, Zile Wang, Wenqing Yu, Donglei Chen, Xuanzhong Sun, Hao Liao, Minpeng Fan, Kai Jiang, Yong Xie, Penguin Zhao, Wayne Xin Song, Ruihua Huang, Fei |
| contents | Large Reasoning Models (LRMs) face two fundamental limitations: excessive token consumption when overanalyzing simple information processing tasks, and inability to access up-to-date knowledge beyond their training data. We introduce MARS (Multi-Agent System for Deep ReSearch), a novel co-evolution framework that jointly optimizes dual cognitive systems through multi-agent reinforcement learning. Unlike prior approaches that employ fixed or independently-trained summarizers, MARS enables System 1 (fast, intuitive processing) and System 2 (deliberate reasoning) to co-adapt through shared trajectory rewards, developing complementary strategies where System 1 learns to distill information specifically useful for System 2's reasoning. We extend Group Relative Policy Optimization (GRPO) for multi-agent settings with three key innovations: (1) decoupled gradient computation ensuring proper credit assignment despite shared rewards, (2) bin-packing optimization for efficient parallel information processing, and (3) advantage-weighted balanced sampling preventing training imbalance. Extensive experiments demonstrate that MARS (8B), trained under a challenging Zero RL setting without any supervised fine-tuning, achieves 8.17% on HLE -- outperforming WebThinker (32B with SFT, 6.87%) and narrowing the gap with proprietary models like Claude 3.7 Sonnet (7.89%) -- while achieving an average gain of 8.9% across 7 knowledge-intensive tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_04935 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning Chen, Guoxin Qiao, Zile Wang, Wenqing Yu, Donglei Chen, Xuanzhong Sun, Hao Liao, Minpeng Fan, Kai Jiang, Yong Xie, Penguin Zhao, Wayne Xin Song, Ruihua Huang, Fei Artificial Intelligence Computation and Language Machine Learning Large Reasoning Models (LRMs) face two fundamental limitations: excessive token consumption when overanalyzing simple information processing tasks, and inability to access up-to-date knowledge beyond their training data. We introduce MARS (Multi-Agent System for Deep ReSearch), a novel co-evolution framework that jointly optimizes dual cognitive systems through multi-agent reinforcement learning. Unlike prior approaches that employ fixed or independently-trained summarizers, MARS enables System 1 (fast, intuitive processing) and System 2 (deliberate reasoning) to co-adapt through shared trajectory rewards, developing complementary strategies where System 1 learns to distill information specifically useful for System 2's reasoning. We extend Group Relative Policy Optimization (GRPO) for multi-agent settings with three key innovations: (1) decoupled gradient computation ensuring proper credit assignment despite shared rewards, (2) bin-packing optimization for efficient parallel information processing, and (3) advantage-weighted balanced sampling preventing training imbalance. Extensive experiments demonstrate that MARS (8B), trained under a challenging Zero RL setting without any supervised fine-tuning, achieves 8.17% on HLE -- outperforming WebThinker (32B with SFT, 6.87%) and narrowing the gap with proprietary models like Claude 3.7 Sonnet (7.89%) -- while achieving an average gain of 8.9% across 7 knowledge-intensive tasks. |
| title | MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning |
| topic | Artificial Intelligence Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.04935 |