MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Guoxin, Qiao, Zile, Wang, Wenqing, Yu, Donglei, Chen, Xuanzhong, Sun, Hao, Liao, Minpeng, Fan, Kai, Jiang, Yong, Xie, Penguin, Zhao, Wayne Xin, Song, Ruihua, Huang, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
by: Qiao, Zile, et al.
Published: (2025)
by: Qiao, Zile, et al.
Published: (2025)
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
by: Chen, Xuanzhong, et al.
Published: (2025)
by: Chen, Xuanzhong, et al.
Published: (2025)
AlphaMath Almost Zero: Process Supervision without Process
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
Step-level Value Preference Optimization for Mathematical Reasoning
by: Chen, Guoxin, et al.
Published: (2024)
by: Chen, Guoxin, et al.
Published: (2024)
ReForm: Reflective Autoformalization with Prospective Bounded Sequence Optimization
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
by: Fu, Biao, et al.
Published: (2025)
by: Fu, Biao, et al.
Published: (2025)
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
by: Chen, Xinjie, et al.
Published: (2025)
by: Chen, Xinjie, et al.
Published: (2025)
Markov Chain of Thought for Efficient Mathematical Reasoning
by: Yang, Wen, et al.
Published: (2024)
by: Yang, Wen, et al.
Published: (2024)
DecoupleSearch: Decouple Planning and Search via Hierarchical Reward Modeling
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
by: Guo, Garvin, et al.
Published: (2026)
by: Guo, Garvin, et al.
Published: (2026)
JURY-RL: Votes Propose, Proofs Dispose for Label-Free RLVR
by: Chen, Xinjie, et al.
Published: (2026)
by: Chen, Xinjie, et al.
Published: (2026)
Toward Autonomous Long-Horizon Engineering for ML Research
by: Chen, Guoxin, et al.
Published: (2026)
by: Chen, Guoxin, et al.
Published: (2026)
MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline
by: Liao, Minpeng, et al.
Published: (2024)
by: Liao, Minpeng, et al.
Published: (2024)
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
by: Zhao, Jiale, et al.
Published: (2026)
by: Zhao, Jiale, et al.
Published: (2026)
Lungfish‐like antero‐labial tooth addition and amphibian‐like enameloid‐enamel transition in the coronoid of a Devonian stem actinopterygian
by: Donglei Chen
Published: (2025)
by: Donglei Chen
Published: (2025)
MARS: A Meta-Adaptive Reinforcement Learning Framework for Risk-Aware Multi-Agent Portfolio Management
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
LLMs Can Achieve High-quality Simultaneous Machine Translation as Efficiently as Offline
by: Fu, Biao, et al.
Published: (2025)
by: Fu, Biao, et al.
Published: (2025)
Deep Learning-Based 3D Seismic Velocity Inversion Under Dual-Domain Sparse Representation
by: Chen, Guoxin, et al.
Published: (2026)
by: Chen, Guoxin, et al.
Published: (2026)
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
MARS-DA: A Hierarchical Reinforcement Learning Framework for Risk-Aware Multi-Agent Bidding in Power Grids
by: Chen, Jiayi, et al.
Published: (2026)
by: Chen, Jiayi, et al.
Published: (2026)
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
by: Chng, Yong Xien, et al.
Published: (2025)
by: Chng, Yong Xien, et al.
Published: (2025)
Initial Model Incorporation for Deep Learning FWI: Pretraining or Denormalization?
by: Chen, Ruihua, et al.
Published: (2025)
by: Chen, Ruihua, et al.
Published: (2025)
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Scaling Agents via Continual Pre-training
by: Su, Liangcai, et al.
Published: (2025)
by: Su, Liangcai, et al.
Published: (2025)
Encouraging Good Processes Without the Need for Good Answers: Reinforcement Learning for LLM Agent Planning
by: Li, Zhiwei, et al.
Published: (2025)
by: Li, Zhiwei, et al.
Published: (2025)
MARTI-MARS$^2$: Scaling Multi-Agent Self-Search via Reinforcement Learning for Code Generation
by: Wang, Shijie, et al.
Published: (2026)
by: Wang, Shijie, et al.
Published: (2026)
Note on explicit construction of conformal generators on the fuzzy sphere
by: Fan, Ruihua
Published: (2024)
by: Fan, Ruihua
Published: (2024)
A Dual-Agent Adversarial Framework for Robust Generalization in Deep Reinforcement Learning
by: Xie, Zhengpeng, et al.
Published: (2025)
by: Xie, Zhengpeng, et al.
Published: (2025)
RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and Treatment
by: Chen, Xuanzhong, et al.
Published: (2024)
by: Chen, Xuanzhong, et al.
Published: (2024)
DynamicBench: Evaluating Real-Time Report Generation in Large Language Models
by: Li, Jingyao, et al.
Published: (2025)
by: Li, Jingyao, et al.
Published: (2025)
Supportiveness-based Knowledge Rewriting for Retrieval-augmented Language Modeling
by: Qiao, Zile, et al.
Published: (2024)
by: Qiao, Zile, et al.
Published: (2024)
AgentFold: Long-Horizon Web Agents with Proactive Context Management
by: Ye, Rui, et al.
Published: (2025)
by: Ye, Rui, et al.
Published: (2025)
The Breakthrough and Confrontation of Mainland Chinese Opera Films in Hong Kong under the Cold War Framework (1953–1957)
by: Du, Jiachen, et al.
Published: (2025)
by: Du, Jiachen, et al.
Published: (2025)
Simple Vertex Algebras Arising From Congruence Subgroups
by: Dai, Xuanzhong, et al.
Published: (2022)
by: Dai, Xuanzhong, et al.
Published: (2022)
The Second Moment of $\mathrm{GL}_4 \times \mathrm{GL}_2$ $L$-functions at Special Points
by: Qi, Zhi, et al.
Published: (2025)
by: Qi, Zhi, et al.
Published: (2025)
Multi-Agent Deep Reinforcement Learning for Energy Efficient Multi-Hop STAR-RIS-Assisted Transmissions
by: Liao, Pei-Hsiang, et al.
Published: (2024)
by: Liao, Pei-Hsiang, et al.
Published: (2024)
Tongyi DeepResearch Technical Report
by: Tongyi DeepResearch Team, et al.
Published: (2025)
by: Tongyi DeepResearch Team, et al.
Published: (2025)
Similar Items
-
IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling
by: Chen, Guoxin, et al.
Published: (2025) -
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
by: Qiao, Zile, et al.
Published: (2025) -
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation
by: Chen, Guoxin, et al.
Published: (2025) -
AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
by: Chen, Xuanzhong, et al.
Published: (2025) -
AlphaMath Almost Zero: Process Supervision without Process
by: Chen, Guoxin, et al.
Published: (2024)