Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914115313729536 |
|---|---|
| author | Fan, Yijia Zhang, Jusheng Yang, Jing Wang, Keze |
| author_facet | Fan, Yijia Zhang, Jusheng Yang, Jing Wang, Keze |
| contents | To combat the prohibitive communication costs of ``free-for-all" multi-agent systems (MAS), we introduce \textbf{Agent-GSPO}, a framework that directly optimizes for token economy using sequence-level reinforcement learning. Agent-GSPO leverages the stable and memory-efficient Group Sequence Policy Optimization (GSPO) algorithm to train agents on a communication-aware reward that explicitly penalizes verbosity. Across seven reasoning benchmarks, Agent-GSPO not only achieves new state-of-the-art performance but does so with a fraction of the token consumption of existing methods. By fostering emergent strategies like ``strategic silence," our approach provides a practical blueprint for developing scalable and economically viable multi-agent systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_22477 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization Fan, Yijia Zhang, Jusheng Yang, Jing Wang, Keze Multiagent Systems Artificial Intelligence To combat the prohibitive communication costs of ``free-for-all" multi-agent systems (MAS), we introduce \textbf{Agent-GSPO}, a framework that directly optimizes for token economy using sequence-level reinforcement learning. Agent-GSPO leverages the stable and memory-efficient Group Sequence Policy Optimization (GSPO) algorithm to train agents on a communication-aware reward that explicitly penalizes verbosity. Across seven reasoning benchmarks, Agent-GSPO not only achieves new state-of-the-art performance but does so with a fraction of the token consumption of existing methods. By fostering emergent strategies like ``strategic silence," our approach provides a practical blueprint for developing scalable and economically viable multi-agent systems. |
| title | Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization |
| topic | Multiagent Systems Artificial Intelligence |
| url | https://arxiv.org/abs/2510.22477 |