Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866908630834479104 |
|---|---|
| author | Turtayev, Rustem Fedorova, Natalia Serikov, Oleg Koldyba, Sergey Avagyan, Lev Volkov, Dmitrii |
| author_facet | Turtayev, Rustem Fedorova, Natalia Serikov, Oleg Koldyba, Sergey Avagyan, Lev Volkov, Dmitrii |
| contents | Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents pursuing unintended or unsafe goals. The bounty received 295 submissions, of which nine were awarded.
This report explains the program's motivation and evaluation criteria, and walks through the nine winning submissions step by step. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_19738 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Misalignment Bounty: Crowdsourcing AI Agent Misbehavior Turtayev, Rustem Fedorova, Natalia Serikov, Oleg Koldyba, Sergey Avagyan, Lev Volkov, Dmitrii Artificial Intelligence Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collected cases of agents pursuing unintended or unsafe goals. The bounty received 295 submissions, of which nine were awarded. This report explains the program's motivation and evaluation criteria, and walks through the nine winning submissions step by step. |
| title | Misalignment Bounty: Crowdsourcing AI Agent Misbehavior |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2510.19738 |