Causal Foundations of Collective Agency

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Jørgensen, Frederik Hytting, Weichwald, Sebastian, Hammond, Lewis
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911637929197568
author Jørgensen, Frederik Hytting
Weichwald, Sebastian
Hammond, Lewis
author_facet Jørgensen, Frederik Hytting
Weichwald, Sebastian
Hammond, Lewis
contents A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining when a group of agents can be viewed as a unified collective agent is a foundational question in the study of interactions and incentives in both biological and artificial systems. We adopt a behavioral perspective in answering this question, ascribing collective agency to a group when viewing the group's joint actions as rational and goal-directed successfully predicts its behavior. We formalize this perspective on collective agency using causal games -- which are causal models of strategic, multi-agent interactions -- and causal abstraction -- which formalizes when a simple, high-level model faithfully captures a more complex, low-level model. We use this framework to solve a puzzle regarding multi-agent incentives in actor-critic models and to make quantitative assessments of the degree of collective agency exhibited by different voting mechanisms. Our framework aims to provide a foundation for theoretical and empirical work to understand, predict, and control emergent collective agents in multi-agent AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00248
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Causal Foundations of Collective Agency
Jørgensen, Frederik Hytting
Weichwald, Sebastian
Hammond, Lewis
Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
A key challenge for the safety of advanced AI systems is the possibility that multiple simpler agents might inadvertently form a collective agent with capabilities and goals distinct from those of any individual. More generally, determining when a group of agents can be viewed as a unified collective agent is a foundational question in the study of interactions and incentives in both biological and artificial systems. We adopt a behavioral perspective in answering this question, ascribing collective agency to a group when viewing the group's joint actions as rational and goal-directed successfully predicts its behavior. We formalize this perspective on collective agency using causal games -- which are causal models of strategic, multi-agent interactions -- and causal abstraction -- which formalizes when a simple, high-level model faithfully captures a more complex, low-level model. We use this framework to solve a puzzle regarding multi-agent incentives in actor-critic models and to make quantitative assessments of the degree of collective agency exhibited by different voting mechanisms. Our framework aims to provide a foundation for theoretical and empirical work to understand, predict, and control emergent collective agents in multi-agent AI systems.
title Causal Foundations of Collective Agency
topic Artificial Intelligence
Computer Science and Game Theory
Multiagent Systems
url https://arxiv.org/abs/2605.00248