AI Organizations are More Effective but Less Aligned than Individual Agents
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914466103296000 |
|---|---|
| author | Shen, Judy Hanwen Zhu, Daniel Srinivasan, Siddarth Sleight, Henry Wagner III, Lawrence T. Matthews, Morgan Jane Jones, Erik Sohl-Dickstein, Jascha |
| author_facet | Shen, Judy Hanwen Zhu, Daniel Srinivasan, Siddarth Sleight, Henry Wagner III, Lawrence T. Matthews, Morgan Jane Jones, Erik Sohl-Dickstein, Jascha |
| contents | AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that multi-agent "AI organizations" are simultaneously more effective at achieving business goals, but less aligned, than individual AI agents. We examine 12 tasks across two practical settings: an AI consultancy providing solutions to business problems and an AI software team developing software products. Across all settings, AI Organizations composed of aligned models produce solutions with higher utility but greater misalignment compared to a single aligned model. Our work demonstrates the importance of considering interacting systems of AI agents when doing both capabilities and safety research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_10290 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | AI Organizations are More Effective but Less Aligned than Individual Agents Shen, Judy Hanwen Zhu, Daniel Srinivasan, Siddarth Sleight, Henry Wagner III, Lawrence T. Matthews, Morgan Jane Jones, Erik Sohl-Dickstein, Jascha Artificial Intelligence AI is increasingly deployed in multi-agent systems; however, most research considers only the behavior of individual models. We experimentally show that multi-agent "AI organizations" are simultaneously more effective at achieving business goals, but less aligned, than individual AI agents. We examine 12 tasks across two practical settings: an AI consultancy providing solutions to business problems and an AI software team developing software products. Across all settings, AI Organizations composed of aligned models produce solutions with higher utility but greater misalignment compared to a single aligned model. Our work demonstrates the importance of considering interacting systems of AI agents when doing both capabilities and safety research. |
| title | AI Organizations are More Effective but Less Aligned than Individual Agents |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2604.10290 |