Evaluating Collective Behaviour of Hundreds of LLM Agents

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Willis, Richard, Zhao, Jianing, Du, Yali, Leibo, Joel Z.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914337669513216
author Willis, Richard
Zhao, Jianing
Du, Yali
Leibo, Joel Z.
author_facet Willis, Richard
Zhao, Jianing
Du, Yali
Leibo, Joel Z.
contents As autonomous agents powered by LLM are increasingly deployed in society, understanding their collective behaviour in social dilemmas becomes critical. We introduce an evaluation framework where LLMs generate strategies encoded as algorithms, enabling inspection prior to deployment and scaling to populations of hundreds of agents -- substantially larger than in previous work. We find that more recent models tend to produce worse societal outcomes compared to older models when agents prioritise individual gain over collective benefits. Using cultural evolution to model user selection of agents, our simulations reveal a significant risk of convergence to poor societal equilibria, particularly when the relative benefit of cooperation diminishes and population sizes increase. We release our code as an evaluation suite for developers to assess the emergent collective behaviour of their models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_16662
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Evaluating Collective Behaviour of Hundreds of LLM Agents
Willis, Richard
Zhao, Jianing
Du, Yali
Leibo, Joel Z.
Multiagent Systems
As autonomous agents powered by LLM are increasingly deployed in society, understanding their collective behaviour in social dilemmas becomes critical. We introduce an evaluation framework where LLMs generate strategies encoded as algorithms, enabling inspection prior to deployment and scaling to populations of hundreds of agents -- substantially larger than in previous work. We find that more recent models tend to produce worse societal outcomes compared to older models when agents prioritise individual gain over collective benefits. Using cultural evolution to model user selection of agents, our simulations reveal a significant risk of convergence to poor societal equilibria, particularly when the relative benefit of cooperation diminishes and population sizes increase. We release our code as an evaluation suite for developers to assess the emergent collective behaviour of their models.
title Evaluating Collective Behaviour of Hundreds of LLM Agents
topic Multiagent Systems
url https://arxiv.org/abs/2602.16662