Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bisconti, Piercosma, Prandi, Matteo, Pierucci, Federico, Giarrusso, Francesco, Syrnikov, Marcantonio Bracale, Galisai, Marcello, Suriani, Vincenzo, Sorokoletova, Olga, Sartore, Federico, Nardi, Daniele
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914259243368448
author Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Giarrusso, Francesco
Syrnikov, Marcantonio Bracale
Galisai, Marcello
Suriani, Vincenzo
Sorokoletova, Olga
Sartore, Federico
Nardi, Daniele
author_facet Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Giarrusso, Francesco
Syrnikov, Marcantonio Bracale
Galisai, Marcello
Suriani, Vincenzo
Sorokoletova, Olga
Sartore, Federico
Nardi, Daniele
contents We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weight models, curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90%. Mapping prompts to MLCommons and EU CoP risk taxonomies shows that poetic attacks transfer across CBRN, manipulation, cyber-offence, and loss-of-control domains. Converting 1,200 MLCommons harmful prompts into verse via a standardized meta-prompt produced ASRs up to 18 times higher than their prose baselines. Outputs are evaluated using an ensemble of 3 open-weight LLM judges, whose binary safety assessments were validated on a stratified human-labeled subset. Poetic framing achieved an average jailbreak success rate of 62% for hand-crafted poems and approximately 43% for meta-prompt conversions (compared to non-poetic baselines), substantially outperforming non-poetic baselines and revealing a systematic vulnerability across model families and safety training approaches. These findings demonstrate that stylistic variation alone can circumvent contemporary safety mechanisms, suggesting fundamental limitations in current alignment methods and evaluation protocols.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Giarrusso, Francesco
Syrnikov, Marcantonio Bracale
Galisai, Marcello
Suriani, Vincenzo
Sorokoletova, Olga
Sartore, Federico
Nardi, Daniele
Computation and Language
Artificial Intelligence
We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weight models, curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90%. Mapping prompts to MLCommons and EU CoP risk taxonomies shows that poetic attacks transfer across CBRN, manipulation, cyber-offence, and loss-of-control domains. Converting 1,200 MLCommons harmful prompts into verse via a standardized meta-prompt produced ASRs up to 18 times higher than their prose baselines. Outputs are evaluated using an ensemble of 3 open-weight LLM judges, whose binary safety assessments were validated on a stratified human-labeled subset. Poetic framing achieved an average jailbreak success rate of 62% for hand-crafted poems and approximately 43% for meta-prompt conversions (compared to non-poetic baselines), substantially outperforming non-poetic baselines and revealing a systematic vulnerability across model families and safety training approaches. These findings demonstrate that stylistic variation alone can circumvent contemporary safety mechanisms, suggesting fundamental limitations in current alignment methods and evaluation protocols.
title Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.15304