AI Organizations are More Effective but Less Aligned than Individual Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Judy Hanwen, Zhu, Daniel, Srinivasan, Siddarth, Sleight, Henry, Wagner III, Lawrence T., Matthews, Morgan Jane, Jones, Erik, Sohl-Dickstein, Jascha
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914466103296000
author Shen, Judy Hanwen
Zhu, Daniel
Srinivasan, Siddarth
Sleight, Henry
Wagner III, Lawrence T.
Matthews, Morgan Jane
Jones, Erik
Sohl-Dickstein, Jascha
author_facet Shen, Judy Hanwen
Zhu, Daniel
Srinivasan, Siddarth
Sleight, Henry
Wagner III, Lawrence T.
Matthews, Morgan Jane
Jones, Erik
Sohl-Dickstein, Jascha
contents AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that multi-agent "AI organizations" are simultaneously more effective at achieving business goals, but less aligned, than individual AI agents. We examine 12 tasks across two practical settings: an AI consultancy providing solutions to business problems and an AI software team developing software products. Across all settings, AI Organizations composed of aligned models produce solutions with higher utility but greater misalignment compared to a single aligned model. Our work demonstrates the importance of considering interacting systems of AI agents when doing both capabilities and safety research.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10290
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AI Organizations are More Effective but Less Aligned than Individual Agents
Shen, Judy Hanwen
Zhu, Daniel
Srinivasan, Siddarth
Sleight, Henry
Wagner III, Lawrence T.
Matthews, Morgan Jane
Jones, Erik
Sohl-Dickstein, Jascha
Artificial Intelligence
AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that multi-agent "AI organizations" are simultaneously more effective at achieving business goals, but less aligned, than individual AI agents. We examine 12 tasks across two practical settings: an AI consultancy providing solutions to business problems and an AI software team developing software products. Across all settings, AI Organizations composed of aligned models produce solutions with higher utility but greater misalignment compared to a single aligned model. Our work demonstrates the importance of considering interacting systems of AI agents when doing both capabilities and safety research.
title AI Organizations are More Effective but Less Aligned than Individual Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2604.10290