Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Yuzhe, Liu, Feiran, Shan, Yi, Huang, Xinyi, Yang, Xin, Zhu, Yueqi, Cheng, Xuxin, Liu, Cao, Zeng, Ke, Zhang, Terry Jingchen, Jiang, Wenyuan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914469870829568
author Zhang, Yuzhe
Liu, Feiran
Shan, Yi
Huang, Xinyi
Yang, Xin
Zhu, Yueqi
Cheng, Xuxin
Liu, Cao
Zeng, Ke
Zhang, Terry Jingchen
Jiang, Wenyuan
author_facet Zhang, Yuzhe
Liu, Feiran
Shan, Yi
Huang, Xinyi
Yang, Xin
Zhu, Yueqi
Cheng, Xuxin
Liu, Cao
Zeng, Ke
Zhang, Terry Jingchen
Jiang, Wenyuan
contents Large language models are increasingly deployed in multi-agent systems to overcome context limitations by distributing information across agents. Yet whether agents can reliably compute with distributed information, rather than merely exchange it, remains an open question. We introduce SILO-BENCH, a role-agnostic benchmark of 30 algorithmic tasks across three communication complexity levels, evaluating 54 configurations over 1,620 experiments. Our experiments expose a fundamental Communication-Reasoning Gap: agents spontaneously form task-appropriate coordination topologies and exchange information actively, yet systematically fail to synthesize distributed state into correct answers. The failure is localized to the reasoning-integration stage where agents often acquire sufficient information but cannot integrate it. This coordination overhead compounds with scale, eventually eliminating parallelization gains entirely. These findings demonstrate that naively scaling agent count cannot circumvent context limitations, and SILO-BENCH provides a foundation for tracking progress toward genuinely collaborative multi-agent systems. The code is available at https://github.com/jwyjohn/acl26-silo-bench .
format Preprint
id arxiv_https___arxiv_org_abs_2603_01045
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
Zhang, Yuzhe
Liu, Feiran
Shan, Yi
Huang, Xinyi
Yang, Xin
Zhu, Yueqi
Cheng, Xuxin
Liu, Cao
Zeng, Ke
Zhang, Terry Jingchen
Jiang, Wenyuan
Multiagent Systems
Artificial Intelligence
Large language models are increasingly deployed in multi-agent systems to overcome context limitations by distributing information across agents. Yet whether agents can reliably compute with distributed information, rather than merely exchange it, remains an open question. We introduce SILO-BENCH, a role-agnostic benchmark of 30 algorithmic tasks across three communication complexity levels, evaluating 54 configurations over 1,620 experiments. Our experiments expose a fundamental Communication-Reasoning Gap: agents spontaneously form task-appropriate coordination topologies and exchange information actively, yet systematically fail to synthesize distributed state into correct answers. The failure is localized to the reasoning-integration stage where agents often acquire sufficient information but cannot integrate it. This coordination overhead compounds with scale, eventually eliminating parallelization gains entirely. These findings demonstrate that naively scaling agent count cannot circumvent context limitations, and SILO-BENCH provides a foundation for tracking progress toward genuinely collaborative multi-agent systems. The code is available at https://github.com/jwyjohn/acl26-silo-bench .
title Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
topic Multiagent Systems
Artificial Intelligence
url https://arxiv.org/abs/2603.01045