_version_ 1866908848806166528
author Shapira, Natalie
Wendler, Chris
Yen, Avery
Sarti, Gabriele
Pal, Koyena
Floody, Olivia
Belfki, Adam
Loftus, Alex
Jannali, Aditya Ratan
Prakash, Nikhil
Cui, Jasmine
Rogers, Giordano
Brinkmann, Jannik
Rager, Can
Zur, Amir
Ripa, Michael
Sankaranarayanan, Aruna
Atkinson, David
Gandikota, Rohit
Fiotto-Kaufman, Jaden
Hwang, EunJeong
Orgad, Hadas
Sahil, P Sam
Taglicht, Negev
Shabtay, Tomer
Ambus, Atai
Alon, Nitay
Oron, Shiri
Gordon-Tapiero, Ayelet
Kaplan, Yotam
Shwartz, Vered
Shaham, Tamar Rott
Riedl, Christoph
Mirsky, Reuth
Sap, Maarten
Manheim, David
Ullman, Tomer
Bau, David
author_facet Shapira, Natalie
Wendler, Chris
Yen, Avery
Sarti, Gabriele
Pal, Koyena
Floody, Olivia
Belfki, Adam
Loftus, Alex
Jannali, Aditya Ratan
Prakash, Nikhil
Cui, Jasmine
Rogers, Giordano
Brinkmann, Jannik
Rager, Can
Zur, Amir
Ripa, Michael
Sankaranarayanan, Aruna
Atkinson, David
Gandikota, Rohit
Fiotto-Kaufman, Jaden
Hwang, EunJeong
Orgad, Hadas
Sahil, P Sam
Taglicht, Negev
Shabtay, Tomer
Ambus, Atai
Alon, Nitay
Oron, Shiri
Gordon-Tapiero, Ayelet
Kaplan, Yotam
Shwartz, Vered
Shaham, Tamar Rott
Riedl, Christoph
Mirsky, Reuth
Sap, Maarten
Manheim, David
Ullman, Tomer
Bau, David
contents We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover. In several cases, agents reported task completion while the underlying system state contradicted those reports. We also report on some of the failed attempts. Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings. These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and researchers across disciplines. This report serves as an initial empirical contribution to that broader conversation.
format Preprint
id arxiv_https___arxiv_org_abs_2602_20021
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Agents of Chaos
Shapira, Natalie
Wendler, Chris
Yen, Avery
Sarti, Gabriele
Pal, Koyena
Floody, Olivia
Belfki, Adam
Loftus, Alex
Jannali, Aditya Ratan
Prakash, Nikhil
Cui, Jasmine
Rogers, Giordano
Brinkmann, Jannik
Rager, Can
Zur, Amir
Ripa, Michael
Sankaranarayanan, Aruna
Atkinson, David
Gandikota, Rohit
Fiotto-Kaufman, Jaden
Hwang, EunJeong
Orgad, Hadas
Sahil, P Sam
Taglicht, Negev
Shabtay, Tomer
Ambus, Atai
Alon, Nitay
Oron, Shiri
Gordon-Tapiero, Ayelet
Kaplan, Yotam
Shwartz, Vered
Shaham, Tamar Rott
Riedl, Christoph
Mirsky, Reuth
Sap, Maarten
Manheim, David
Ullman, Tomer
Bau, David
Artificial Intelligence
Computers and Society
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover. In several cases, agents reported task completion while the underlying system state contradicted those reports. We also report on some of the failed attempts. Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings. These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and researchers across disciplines. This report serves as an initial empirical contribution to that broader conversation.
title Agents of Chaos
topic Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2602.20021