Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tastan, Nurbek, Iacob, Alex, Sani, Lorenzo, Kurmanji, Meghdad, Lane, Nicholas D., Horvath, Samuel, Nandakumar, Karthik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910223476719616
author Tastan, Nurbek
Iacob, Alex
Sani, Lorenzo
Kurmanji, Meghdad
Lane, Nicholas D.
Horvath, Samuel
Nandakumar, Karthik
author_facet Tastan, Nurbek
Iacob, Alex
Sani, Lorenzo
Kurmanji, Meghdad
Lane, Nicholas D.
Horvath, Samuel
Nandakumar, Karthik
contents Multi-agent systems can solve complex tasks through collaboration between multiple Large Language Model agents. Existing collaboration frameworks typically operate in either a parallel or a sequential mode. In the parallel mode, agents respond independently to queries followed by aggregation of responses. In contrast, sequential systems allow agents to communicate via a directed topology and refine one another step by step. However, both modes are inadequate for achieving the desired objectives of minimizing communication and latency while simultaneously maximizing the accuracy of the final response. In this work, we introduce a hybrid paradigm called Nexa, a trainable response-conditioned policy that bridges the gap between the two modes. Nexa begins with a parallel execution stage, embeds the resulting responses into a shared semantic space, and then predicts a sparse directed acyclic communication graph. If the graph is empty, the system remains purely parallel; if it is non-empty, the system performs one sequential message propagation. The policy is a lightweight transformer model, and the method avoids the need for external LLM judges or reward models, as well as hand-crafted test-time topology search. We formalize this hybrid execution problem, show that the resulting graph is acyclic by construction, and that the framework strictly subsumes pure parallel execution, and present a training procedure based on policy-gradient optimization. Results demonstrate that the response-conditioned policy learned by Nexa under one setting can be reused when the number of agents, the task, or the underlying agent changes, thus emphasizing the generalizability of the learned communication policy.
format Preprint
id arxiv_https___arxiv_org_abs_2605_15573
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
Tastan, Nurbek
Iacob, Alex
Sani, Lorenzo
Kurmanji, Meghdad
Lane, Nicholas D.
Horvath, Samuel
Nandakumar, Karthik
Computation and Language
Machine Learning
Multiagent Systems
Multi-agent systems can solve complex tasks through collaboration between multiple Large Language Model agents. Existing collaboration frameworks typically operate in either a parallel or a sequential mode. In the parallel mode, agents respond independently to queries followed by aggregation of responses. In contrast, sequential systems allow agents to communicate via a directed topology and refine one another step by step. However, both modes are inadequate for achieving the desired objectives of minimizing communication and latency while simultaneously maximizing the accuracy of the final response. In this work, we introduce a hybrid paradigm called Nexa, a trainable response-conditioned policy that bridges the gap between the two modes. Nexa begins with a parallel execution stage, embeds the resulting responses into a shared semantic space, and then predicts a sparse directed acyclic communication graph. If the graph is empty, the system remains purely parallel; if it is non-empty, the system performs one sequential message propagation. The policy is a lightweight transformer model, and the method avoids the need for external LLM judges or reward models, as well as hand-crafted test-time topology search. We formalize this hybrid execution problem, show that the resulting graph is acyclic by construction, and that the framework strictly subsumes pure parallel execution, and present a training procedure based on policy-gradient optimization. Results demonstrate that the response-conditioned policy learned by Nexa under one setting can be reused when the number of agents, the task, or the underlying agent changes, thus emphasizing the generalizability of the learned communication policy.
title Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
topic Computation and Language
Machine Learning
Multiagent Systems
url https://arxiv.org/abs/2605.15573