Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Berg, Cameron, Schneider, Susan L., Bailey, Mark M.
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913099262459904
author Berg, Cameron
Schneider, Susan L.
Bailey, Mark M.
author_facet Berg, Cameron
Schneider, Susan L.
Bailey, Mark M.
contents Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. However, observing agent behavior alone is often insufficient to distinguish genuine informational coupling from spurious similarity, as consequential coalitions may form at the level of internal representations before any overt behavioral change is apparent. Here, we introduce a practical method for detecting coalition structure from the internal neural representations of multi-agent systems. The approach constructs a pairwise mutual-information graph from the hidden states of agents and applies spectral partitioning to identify the most salient coalition boundary. We validate this method in two domains. First, in multi-agent reinforcement learning environments, the method successfully recovers programmed hierarchical and dynamic coalition structures and correctly rejects false positives arising from behavioral coordination without informational coupling. Second, using a large language model, the method identifies coalition structures implied by descriptive prompts, tracks dynamic team reassignments, and reveals a representational hierarchy where explicit labels dominate over conflicting interaction patterns. Across both settings, the recovered partition reveals subgroup organization that a scalar cross-agent mutual-information measure cannot distinguish. The results demonstrate that analyzing hidden-state mutual information through spectral partitioning provides a scalable diagnostic for identifying representational coalitions, offering a valuable tool for monitoring emergent structure in distributed AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06696
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations
Berg, Cameron
Schneider, Susan L.
Bailey, Mark M.
Artificial Intelligence
Machine Learning
Multiagent Systems
Primary: 68T42, Secondary: 93A16, 68T07, 94A17, 05C50
Collections of interacting AI agents can form coalitions, creating emergent group-level organization that is critical for AI safety and alignment. However, observing agent behavior alone is often insufficient to distinguish genuine informational coupling from spurious similarity, as consequential coalitions may form at the level of internal representations before any overt behavioral change is apparent. Here, we introduce a practical method for detecting coalition structure from the internal neural representations of multi-agent systems. The approach constructs a pairwise mutual-information graph from the hidden states of agents and applies spectral partitioning to identify the most salient coalition boundary. We validate this method in two domains. First, in multi-agent reinforcement learning environments, the method successfully recovers programmed hierarchical and dynamic coalition structures and correctly rejects false positives arising from behavioral coordination without informational coupling. Second, using a large language model, the method identifies coalition structures implied by descriptive prompts, tracks dynamic team reassignments, and reveals a representational hierarchy where explicit labels dominate over conflicting interaction patterns. Across both settings, the recovered partition reveals subgroup organization that a scalar cross-agent mutual-information measure cannot distinguish. The results demonstrate that analyzing hidden-state mutual information through spectral partitioning provides a scalable diagnostic for identifying representational coalitions, offering a valuable tool for monitoring emergent structure in distributed AI systems.
title Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations
topic Artificial Intelligence
Machine Learning
Multiagent Systems
Primary: 68T42, Secondary: 93A16, 68T07, 94A17, 05C50
url https://arxiv.org/abs/2605.06696