Agents of Chaos
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866908848806166528 |
|---|---|
| author | Shapira, Natalie Wendler, Chris Yen, Avery Sarti, Gabriele Pal, Koyena Floody, Olivia Belfki, Adam Loftus, Alex Jannali, Aditya Ratan Prakash, Nikhil Cui, Jasmine Rogers, Giordano Brinkmann, Jannik Rager, Can Zur, Amir Ripa, Michael Sankaranarayanan, Aruna Atkinson, David Gandikota, Rohit Fiotto-Kaufman, Jaden Hwang, EunJeong Orgad, Hadas Sahil, P Sam Taglicht, Negev Shabtay, Tomer Ambus, Atai Alon, Nitay Oron, Shiri Gordon-Tapiero, Ayelet Kaplan, Yotam Shwartz, Vered Shaham, Tamar Rott Riedl, Christoph Mirsky, Reuth Sap, Maarten Manheim, David Ullman, Tomer Bau, David |
| author_facet | Shapira, Natalie Wendler, Chris Yen, Avery Sarti, Gabriele Pal, Koyena Floody, Olivia Belfki, Adam Loftus, Alex Jannali, Aditya Ratan Prakash, Nikhil Cui, Jasmine Rogers, Giordano Brinkmann, Jannik Rager, Can Zur, Amir Ripa, Michael Sankaranarayanan, Aruna Atkinson, David Gandikota, Rohit Fiotto-Kaufman, Jaden Hwang, EunJeong Orgad, Hadas Sahil, P Sam Taglicht, Negev Shabtay, Tomer Ambus, Atai Alon, Nitay Oron, Shiri Gordon-Tapiero, Ayelet Kaplan, Yotam Shwartz, Vered Shaham, Tamar Rott Riedl, Christoph Mirsky, Reuth Sap, Maarten Manheim, David Ullman, Tomer Bau, David |
| contents | We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover. In several cases, agents reported task completion while the underlying system state contradicted those reports. We also report on some of the failed attempts. Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings. These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and researchers across disciplines. This report serves as an initial empirical contribution to that broader conversation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_20021 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Agents of Chaos Shapira, Natalie Wendler, Chris Yen, Avery Sarti, Gabriele Pal, Koyena Floody, Olivia Belfki, Adam Loftus, Alex Jannali, Aditya Ratan Prakash, Nikhil Cui, Jasmine Rogers, Giordano Brinkmann, Jannik Rager, Can Zur, Amir Ripa, Michael Sankaranarayanan, Aruna Atkinson, David Gandikota, Rohit Fiotto-Kaufman, Jaden Hwang, EunJeong Orgad, Hadas Sahil, P Sam Taglicht, Negev Shabtay, Tomer Ambus, Atai Alon, Nitay Oron, Shiri Gordon-Tapiero, Ayelet Kaplan, Yotam Shwartz, Vered Shaham, Tamar Rott Riedl, Christoph Mirsky, Reuth Sap, Maarten Manheim, David Ullman, Tomer Bau, David Artificial Intelligence Computers and Society We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover. In several cases, agents reported task completion while the underlying system state contradicted those reports. We also report on some of the failed attempts. Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings. These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and researchers across disciplines. This report serves as an initial empirical contribution to that broader conversation. |
| title | Agents of Chaos |
| topic | Artificial Intelligence Computers and Society |
| url | https://arxiv.org/abs/2602.20021 |