Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Yijia, Zhang, Jusheng, Yang, Jing, Wang, Keze
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914115313729536
author Fan, Yijia
Zhang, Jusheng
Yang, Jing
Wang, Keze
author_facet Fan, Yijia
Zhang, Jusheng
Yang, Jing
Wang, Keze
contents To combat the prohibitive communication costs of ``free-for-all" multi-agent systems (MAS), we introduce \textbf{Agent-GSPO}, a framework that directly optimizes for token economy using sequence-level reinforcement learning. Agent-GSPO leverages the stable and memory-efficient Group Sequence Policy Optimization (GSPO) algorithm to train agents on a communication-aware reward that explicitly penalizes verbosity. Across seven reasoning benchmarks, Agent-GSPO not only achieves new state-of-the-art performance but does so with a fraction of the token consumption of existing methods. By fostering emergent strategies like ``strategic silence," our approach provides a practical blueprint for developing scalable and economically viable multi-agent systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization
Fan, Yijia
Zhang, Jusheng
Yang, Jing
Wang, Keze
Multiagent Systems
Artificial Intelligence
To combat the prohibitive communication costs of ``free-for-all" multi-agent systems (MAS), we introduce \textbf{Agent-GSPO}, a framework that directly optimizes for token economy using sequence-level reinforcement learning. Agent-GSPO leverages the stable and memory-efficient Group Sequence Policy Optimization (GSPO) algorithm to train agents on a communication-aware reward that explicitly penalizes verbosity. Across seven reasoning benchmarks, Agent-GSPO not only achieves new state-of-the-art performance but does so with a fraction of the token consumption of existing methods. By fostering emergent strategies like ``strategic silence," our approach provides a practical blueprint for developing scalable and economically viable multi-agent systems.
title Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization
topic Multiagent Systems
Artificial Intelligence
url https://arxiv.org/abs/2510.22477