Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911218255527936 |
|---|---|
| author | Bertollo, Giacomo Bodemir, Naz Burgess, Jonah |
| author_facet | Bertollo, Giacomo Bodemir, Naz Burgess, Jonah |
| contents | Analyzing 500 CTF participants, this paper shows that while participants readily bypassed simple AI guardrails using common techniques, layered multi-step defenses still posed significant challenges, offering concrete insights for building safer AI systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_16005 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers Bertollo, Giacomo Bodemir, Naz Burgess, Jonah Cryptography and Security Artificial Intelligence Analyzing 500 CTF participants, this paper shows that while participants readily bypassed simple AI guardrails using common techniques, layered multi-step defenses still posed significant challenges, offering concrete insights for building safer AI systems. |
| title | Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2510.16005 |