MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Kuo, Yang, Xingjie, Yu, Linhui, Xu, Qing, Fang, Yan, Wang, Xu, Zhou, Zhengyang, Wang, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912426639753216
author Yang, Kuo
Yang, Xingjie
Yu, Linhui
Xu, Qing
Fang, Yan
Wang, Xu
Zhou, Zhengyang
Wang, Yang
author_facet Yang, Kuo
Yang, Xingjie
Yu, Linhui
Xu, Qing
Fang, Yan
Wang, Xu
Zhou, Zhengyang
Wang, Yang
contents Large Language Model (LLM)-driven Multi-agent systems (Mas) have recently emerged as a powerful paradigm for tackling complex real-world tasks. However, existing Mas construction methods typically rely on manually crafted interaction mechanisms or heuristic rules, introducing human biases and constraining the autonomous ability. Even with recent advances in adaptive Mas construction, existing systems largely remain within the paradigm of semi-autonomous patterns. In this work, we propose MasHost, a Reinforcement Learning (RL)-based framework for autonomous and query-adaptive Mas design. By formulating Mas construction as a graph search problem, our proposed MasHost jointly samples agent roles and their interactions through a unified probabilistic sampling mechanism. Beyond the accuracy and efficiency objectives pursued in prior works, we introduce component rationality as an additional and novel design principle in Mas. To achieve this multi-objective optimization, we propose Hierarchical Relative Policy Optimization (HRPO), a novel RL strategy that collaboratively integrates group-relative advantages and action-wise rewards. To our knowledge, our proposed MasHost is the first RL-driven framework for autonomous Mas graph construction. Extensive experiments on six benchmarks demonstrate that MasHost consistently outperforms most competitive baselines, validating its effectiveness, efficiency, and structure rationality.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08507
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning
Yang, Kuo
Yang, Xingjie
Yu, Linhui
Xu, Qing
Fang, Yan
Wang, Xu
Zhou, Zhengyang
Wang, Yang
Multiagent Systems
Artificial Intelligence
Machine Learning
Large Language Model (LLM)-driven Multi-agent systems (Mas) have recently emerged as a powerful paradigm for tackling complex real-world tasks. However, existing Mas construction methods typically rely on manually crafted interaction mechanisms or heuristic rules, introducing human biases and constraining the autonomous ability. Even with recent advances in adaptive Mas construction, existing systems largely remain within the paradigm of semi-autonomous patterns. In this work, we propose MasHost, a Reinforcement Learning (RL)-based framework for autonomous and query-adaptive Mas design. By formulating Mas construction as a graph search problem, our proposed MasHost jointly samples agent roles and their interactions through a unified probabilistic sampling mechanism. Beyond the accuracy and efficiency objectives pursued in prior works, we introduce component rationality as an additional and novel design principle in Mas. To achieve this multi-objective optimization, we propose Hierarchical Relative Policy Optimization (HRPO), a novel RL strategy that collaboratively integrates group-relative advantages and action-wise rewards. To our knowledge, our proposed MasHost is the first RL-driven framework for autonomous Mas graph construction. Extensive experiments on six benchmarks demonstrate that MasHost consistently outperforms most competitive baselines, validating its effectiveness, efficiency, and structure rationality.
title MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning
topic Multiagent Systems
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.08507