Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Bertollo, Giacomo, Bodemir, Naz, Burgess, Jonah
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911218255527936
author Bertollo, Giacomo
Bodemir, Naz
Burgess, Jonah
author_facet Bertollo, Giacomo
Bodemir, Naz
Burgess, Jonah
contents Analyzing 500 CTF participants, this paper shows that while participants readily bypassed simple AI guardrails using common techniques, layered multi-step defenses still posed significant challenges, offering concrete insights for building safer AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_16005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
Bertollo, Giacomo
Bodemir, Naz
Burgess, Jonah
Cryptography and Security
Artificial Intelligence
Analyzing 500 CTF participants, this paper shows that while participants readily bypassed simple AI guardrails using common techniques, layered multi-step defenses still posed significant challenges, offering concrete insights for building safer AI systems.
title Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.16005