Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Ye, Liu, Heming, Jin, Haibo, Yuan, Xiaopeng, Kuang, Peng, Wang, Haohan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913057987362816
author Yu, Ye
Liu, Heming
Jin, Haibo
Yuan, Xiaopeng
Kuang, Peng
Wang, Haohan
author_facet Yu, Ye
Liu, Heming
Jin, Haibo
Yuan, Xiaopeng
Kuang, Peng
Wang, Haohan
contents Multi-agent systems built on large language models have shown strong performance on complex reasoning tasks, yet most work focuses on agent roles and orchestration while treating inter-agent communication as a fixed interface. Latent communication through internal representations such as key-value caches offers a promising alternative to text-based protocols, but existing approaches do not jointly optimize communication with multi-agent reasoning. Therefore we propose DiffMAS, a training framework that treats latent communication as a learnable component of multi-agent systems. DiffMAS performs parameter-efficient supervised training over multi-agent latent trajectories, enabling agents to jointly learn how information should be encoded and interpreted across interactions. Experiments on mathematical reasoning, scientific QA, code generation, and commonsense benchmarks show that DiffMAS consistently improves reasoning accuracy and decoding stability over single-agent inference, text-based multi-agent systems, and prior latent communication methods, achieving 26.7% on AIME24, 20.2% on GPQA-Diamond, and consistent gains across reasoning benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21794
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
Yu, Ye
Liu, Heming
Jin, Haibo
Yuan, Xiaopeng
Kuang, Peng
Wang, Haohan
Artificial Intelligence
Computation and Language
Multiagent Systems
Multi-agent systems built on large language models have shown strong performance on complex reasoning tasks, yet most work focuses on agent roles and orchestration while treating inter-agent communication as a fixed interface. Latent communication through internal representations such as key-value caches offers a promising alternative to text-based protocols, but existing approaches do not jointly optimize communication with multi-agent reasoning. Therefore we propose DiffMAS, a training framework that treats latent communication as a learnable component of multi-agent systems. DiffMAS performs parameter-efficient supervised training over multi-agent latent trajectories, enabling agents to jointly learn how information should be encoded and interpreted across interactions. Experiments on mathematical reasoning, scientific QA, code generation, and commonsense benchmarks show that DiffMAS consistently improves reasoning accuracy and decoding stability over single-agent inference, text-based multi-agent systems, and prior latent communication methods, achieving 26.7% on AIME24, 20.2% on GPQA-Diamond, and consistent gains across reasoning benchmarks.
title Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
topic Artificial Intelligence
Computation and Language
Multiagent Systems
url https://arxiv.org/abs/2604.21794