CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xue, Xiangyuan, Zhou, Yifan, Zhang, Guibin, Zhang, Zaibin, Li, Yijiang, Zhang, Chen, Yin, Zhenfei, Torr, Philip, Ouyang, Wanli, Bai, Lei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915785552691200
author Xue, Xiangyuan
Zhou, Yifan
Zhang, Guibin
Zhang, Zaibin
Li, Yijiang
Zhang, Chen
Yin, Zhenfei
Torr, Philip
Ouyang, Wanli
Bai, Lei
author_facet Xue, Xiangyuan
Zhou, Yifan
Zhang, Guibin
Zhang, Zaibin
Li, Yijiang
Zhang, Chen
Yin, Zhenfei
Torr, Philip
Ouyang, Wanli
Bai, Lei
contents Self-evolution is a central research topic in enabling large language model (LLM)-based agents to continually improve their capabilities after pretraining. Recent research has witnessed a transition from reinforcement learning (RL)-free to RL-based methods. Current RL-based methods either rely on dense external reward signals or extract intrinsic reward signals from LLMs themselves. However, these approaches diverge from the self-evolution mechanisms observed in human intelligence, where individuals learn and improve through mutual discussion and collaboration. In this work, we introduce Co-Evolving Multi-Agent Systems (CoMAS), a novel framework that enables agents to improve autonomously by learning from inter-agent interactions without external supervision. CoMAS generates intrinsic rewards from rich discussion dynamics, employs an LLM-as-a-judge mechanism to formulate these rewards, and optimizes each agent's policy through RL, thereby enabling decentralized and scalable co-evolution. Experimental results demonstrate that CoMAS consistently outperforms untrained agents and achieves state-of-the-art performance across most evaluation settings. Ablation studies confirm the necessity of interaction-based reward signals and reveal promising scalability as the number and diversity of agents increase. These findings establish CoMAS as a novel and effective paradigm for self-evolution in LLM-based agents.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08529
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
Xue, Xiangyuan
Zhou, Yifan
Zhang, Guibin
Zhang, Zaibin
Li, Yijiang
Zhang, Chen
Yin, Zhenfei
Torr, Philip
Ouyang, Wanli
Bai, Lei
Computation and Language
Artificial Intelligence
Self-evolution is a central research topic in enabling large language model (LLM)-based agents to continually improve their capabilities after pretraining. Recent research has witnessed a transition from reinforcement learning (RL)-free to RL-based methods. Current RL-based methods either rely on dense external reward signals or extract intrinsic reward signals from LLMs themselves. However, these approaches diverge from the self-evolution mechanisms observed in human intelligence, where individuals learn and improve through mutual discussion and collaboration. In this work, we introduce Co-Evolving Multi-Agent Systems (CoMAS), a novel framework that enables agents to improve autonomously by learning from inter-agent interactions without external supervision. CoMAS generates intrinsic rewards from rich discussion dynamics, employs an LLM-as-a-judge mechanism to formulate these rewards, and optimizes each agent's policy through RL, thereby enabling decentralized and scalable co-evolution. Experimental results demonstrate that CoMAS consistently outperforms untrained agents and achieves state-of-the-art performance across most evaluation settings. Ablation studies confirm the necessity of interaction-based reward signals and reveal promising scalability as the number and diversity of agents increase. These findings establish CoMAS as a novel and effective paradigm for self-evolution in LLM-based agents.
title CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.08529