Metaphor Is Not All Attention Needs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sorokoletova, Olga, Giarrusso, Francesco, De Luca, Giacomo, Bisconti, Piercosma, Prandi, Matteo, Pierucci, Federico, Galisai, Marcello, Suriani, Vincenzo, Nardi, Daniele
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909036341886976
author Sorokoletova, Olga
Giarrusso, Francesco
De Luca, Giacomo
Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Galisai, Marcello
Suriani, Vincenzo
Nardi, Daniele
author_facet Sorokoletova, Olga
Giarrusso, Francesco
De Luca, Giacomo
Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Galisai, Marcello
Suriani, Vincenzo
Nardi, Daniele
contents Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Although post-training aims to make models robust against many jailbreak strategies, recent evidence shows that stylistic reformulations, such as poetic transformation, can still bypass safety mechanisms with alarming effectiveness. This raises a central question: why do literary jailbreaks succeed? In this work, we investigate whether their effectiveness depends on specific poetic devices, on a failure to recognize literary formatting, or on deeper changes in how models process stylistically irregular prompts. We address this problem through an interpretability analysis of attention patterns. We perform input-level ablation studies to assess the contribution of individual and combinations of poetic devices; construct an interpretable vector representation of attention maps; cluster these representations and train linear probes to predict safety outcomes and literary format. Our results show that models distinguish poetic from prose formats with high accuracy, yet struggle to predict jailbreak success within each format. Clustering further reveals clear separation by literary format, but not by safety label. These findings indicate that jailbreak success is not caused by a failure to recognize poetic formatting; rather, poetic prompts induce distinct processing patterns that remain largely independent of harmful-content detection. Overall, literary jailbreaks appear to misalign large language models not through any single poetic device, but through accumulated stylistic irregularities that alter prompt processing and avoid lexical triggers considered during post-training. This suggests that robustness requires safety mechanisms that account for style-induced shifts in model behavior. We use Qwen3-14B as a representative open-weight case study.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12128
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Metaphor Is Not All Attention Needs
Sorokoletova, Olga
Giarrusso, Francesco
De Luca, Giacomo
Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Galisai, Marcello
Suriani, Vincenzo
Nardi, Daniele
Computation and Language
Computers and Society
Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Although post-training aims to make models robust against many jailbreak strategies, recent evidence shows that stylistic reformulations, such as poetic transformation, can still bypass safety mechanisms with alarming effectiveness. This raises a central question: why do literary jailbreaks succeed? In this work, we investigate whether their effectiveness depends on specific poetic devices, on a failure to recognize literary formatting, or on deeper changes in how models process stylistically irregular prompts. We address this problem through an interpretability analysis of attention patterns. We perform input-level ablation studies to assess the contribution of individual and combinations of poetic devices; construct an interpretable vector representation of attention maps; cluster these representations and train linear probes to predict safety outcomes and literary format. Our results show that models distinguish poetic from prose formats with high accuracy, yet struggle to predict jailbreak success within each format. Clustering further reveals clear separation by literary format, but not by safety label. These findings indicate that jailbreak success is not caused by a failure to recognize poetic formatting; rather, poetic prompts induce distinct processing patterns that remain largely independent of harmful-content detection. Overall, literary jailbreaks appear to misalign large language models not through any single poetic device, but through accumulated stylistic irregularities that alter prompt processing and avoid lexical triggers considered during post-training. This suggests that robustness requires safety mechanisms that account for style-induced shifts in model behavior. We use Qwen3-14B as a representative open-weight case study.
title Metaphor Is Not All Attention Needs
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2605.12128