Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Weizhen, Lin, Jianbo, Jiang, Zhuosong, Cao, Jingyi, Liu, Xinpeng, Zhang, Jiayu, Huang, Zhenqiang, Chen, Qianben, Sun, Weichen, Wang, Qiexiang, Lu, Hongxuan, Qin, Tianrui, Zhu, Chenghao, Yao, Yi, Fan, Shuying, Li, Xiaowan, Wang, Tiannan, Liu, Pai, Zhu, King, Zhu, He, Shi, Dingfeng, Wang, Piaohong, Guan, Yeyi, Tang, Xiangru, Liu, Minghao, Jiang, Yuchen Eleanor, Yang, Jian, Liu, Jiaheng, Zhang, Ge, Zhou, Wangchunshu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912541675880448
author Li, Weizhen
Lin, Jianbo
Jiang, Zhuosong
Cao, Jingyi
Liu, Xinpeng
Zhang, Jiayu
Huang, Zhenqiang
Chen, Qianben
Sun, Weichen
Wang, Qiexiang
Lu, Hongxuan
Qin, Tianrui
Zhu, Chenghao
Yao, Yi
Fan, Shuying
Li, Xiaowan
Wang, Tiannan
Liu, Pai
Zhu, King
Zhu, He
Shi, Dingfeng
Wang, Piaohong
Guan, Yeyi
Tang, Xiangru
Liu, Minghao
Jiang, Yuchen Eleanor
Yang, Jian
Liu, Jiaheng
Zhang, Ge
Zhou, Wangchunshu
author_facet Li, Weizhen
Lin, Jianbo
Jiang, Zhuosong
Cao, Jingyi
Liu, Xinpeng
Zhang, Jiayu
Huang, Zhenqiang
Chen, Qianben
Sun, Weichen
Wang, Qiexiang
Lu, Hongxuan
Qin, Tianrui
Zhu, Chenghao
Yao, Yi
Fan, Shuying
Li, Xiaowan
Wang, Tiannan
Liu, Pai
Zhu, King
Zhu, He
Shi, Dingfeng
Wang, Piaohong
Guan, Yeyi
Tang, Xiangru
Liu, Minghao
Jiang, Yuchen Eleanor
Yang, Jian
Liu, Jiaheng
Zhang, Ge
Zhou, Wangchunshu
contents Recent advances in large language models (LLMs) and multi-agent systems have demonstrated remarkable capabilities in complex problem-solving tasks such as deep research, vibe coding, and mathematical reasoning. However, most existing multi-agent systems are built upon manual prompt/workflow engineering with sophisticated agent frameworks, making them computationally inefficient, less capable, and can not benefit from data-centric learning. In this work, we introduce Chain-of-Agents (CoA), a novel paradigm of LLM reasoning that enables native end-to-end complex problem-solving in the same way as a multi-agent system (i.e., multi-turn problem solving with multiple tools and multiple agents) within one model. In chain-of-agents problem-solving, the model dynamically activates different tool agents and role-playing agents to simulate multi-agent collaboration in an end-to-end fashion. To elicit end-to-end chain-of-agents problem-solving abilities in LLMs, we introduce a multi-agent distillation framework to distill state-of-the-art multi-agent systems into chain-of-agents trajectories for agentic supervised fine-tuning. We then use agentic reinforcement learning on verifiable agentic tasks to further improve the models' capabilities on chain-of-agents problem solving. We call the resulting models Agent Foundation Models (AFMs). Our empirical studies demonstrate that AFM establishes new state-of-the-art performance across diverse benchmarks in both web agent and code agent settings. We make the entire research, including the model weights, code for training and evaluation, and the training data, fully open-sourced, which offers a solid starting point for future research on agent models and agentic RL.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Li, Weizhen
Lin, Jianbo
Jiang, Zhuosong
Cao, Jingyi
Liu, Xinpeng
Zhang, Jiayu
Huang, Zhenqiang
Chen, Qianben
Sun, Weichen
Wang, Qiexiang
Lu, Hongxuan
Qin, Tianrui
Zhu, Chenghao
Yao, Yi
Fan, Shuying
Li, Xiaowan
Wang, Tiannan
Liu, Pai
Zhu, King
Zhu, He
Shi, Dingfeng
Wang, Piaohong
Guan, Yeyi
Tang, Xiangru
Liu, Minghao
Jiang, Yuchen Eleanor
Yang, Jian
Liu, Jiaheng
Zhang, Ge
Zhou, Wangchunshu
Artificial Intelligence
Computation and Language
Recent advances in large language models (LLMs) and multi-agent systems have demonstrated remarkable capabilities in complex problem-solving tasks such as deep research, vibe coding, and mathematical reasoning. However, most existing multi-agent systems are built upon manual prompt/workflow engineering with sophisticated agent frameworks, making them computationally inefficient, less capable, and can not benefit from data-centric learning. In this work, we introduce Chain-of-Agents (CoA), a novel paradigm of LLM reasoning that enables native end-to-end complex problem-solving in the same way as a multi-agent system (i.e., multi-turn problem solving with multiple tools and multiple agents) within one model. In chain-of-agents problem-solving, the model dynamically activates different tool agents and role-playing agents to simulate multi-agent collaboration in an end-to-end fashion. To elicit end-to-end chain-of-agents problem-solving abilities in LLMs, we introduce a multi-agent distillation framework to distill state-of-the-art multi-agent systems into chain-of-agents trajectories for agentic supervised fine-tuning. We then use agentic reinforcement learning on verifiable agentic tasks to further improve the models' capabilities on chain-of-agents problem solving. We call the resulting models Agent Foundation Models (AFMs). Our empirical studies demonstrate that AFM establishes new state-of-the-art performance across diverse benchmarks in both web agent and code agent settings. We make the entire research, including the model weights, code for training and evaluation, and the training data, fully open-sourced, which offers a solid starting point for future research on agent models and agentic RL.
title Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.13167