Can AI Agents Agree?

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Berdoz, Frédéric, Rugli, Leonardo, Wattenhofer, Roger
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910050202681344
author Berdoz, Frédéric
Rugli, Leonardo
Wattenhofer, Roger
author_facet Berdoz, Frédéric
Rugli, Leonardo
Wattenhofer, Roger
contents Large language models are increasingly deployed as cooperating agents, yet their behavior in adversarial consensus settings has not been systematically studied. We evaluate LLM-based agents on a Byzantine consensus game over scalar values using a synchronous all-to-all simulation. We test consensus in a no-stake setting where agents have no preferences over the final value, so evaluation focuses on agreement rather than value optimality. Across hundreds of simulations spanning model sizes, group sizes, and Byzantine fractions, we find that valid agreement is not reliable even in benign settings and degrades as group size grows. Introducing a small number of Byzantine agents further reduces success. Failures are dominated by loss of liveness, such as timeouts and stalled convergence, rather than subtle value corruption. Overall, the results suggest that reliable agreement is not yet a dependable emergent capability of current LLM-agent groups even in no-stake settings, raising caution for deployments that rely on robust coordination.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01213
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can AI Agents Agree?
Berdoz, Frédéric
Rugli, Leonardo
Wattenhofer, Roger
Multiagent Systems
Machine Learning
Large language models are increasingly deployed as cooperating agents, yet their behavior in adversarial consensus settings has not been systematically studied. We evaluate LLM-based agents on a Byzantine consensus game over scalar values using a synchronous all-to-all simulation. We test consensus in a no-stake setting where agents have no preferences over the final value, so evaluation focuses on agreement rather than value optimality. Across hundreds of simulations spanning model sizes, group sizes, and Byzantine fractions, we find that valid agreement is not reliable even in benign settings and degrades as group size grows. Introducing a small number of Byzantine agents further reduces success. Failures are dominated by loss of liveness, such as timeouts and stalled convergence, rather than subtle value corruption. Overall, the results suggest that reliable agreement is not yet a dependable emergent capability of current LLM-agent groups even in no-stake settings, raising caution for deployments that rely on robust coordination.
title Can AI Agents Agree?
topic Multiagent Systems
Machine Learning
url https://arxiv.org/abs/2603.01213