Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Noughabi, Havva Alizadeh, Serbanescu, Julien, Zarrinkalam, Fattane, Dehghantanha, Ali
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908611383394304
author Noughabi, Havva Alizadeh
Serbanescu, Julien
Zarrinkalam, Fattane
Dehghantanha, Ali
author_facet Noughabi, Havva Alizadeh
Serbanescu, Julien
Zarrinkalam, Fattane
Dehghantanha, Ali
contents Despite recent advances, Large Language Models remain vulnerable to jailbreak attacks that bypass alignment safeguards and elicit harmful outputs. While prior research has proposed various attack strategies differing in human readability and transferability, little attention has been paid to the linguistic and psychological mechanisms that may influence a model's susceptibility to such attacks. In this paper, we examine an interdisciplinary line of research that leverages foundational theories of persuasion from the social sciences to craft adversarial prompts capable of circumventing alignment constraints in LLMs. Drawing on well-established persuasive strategies, we hypothesize that LLMs, having been trained on large-scale human-generated text, may respond more compliantly to prompts with persuasive structures. Furthermore, we investigate whether LLMs themselves exhibit distinct persuasive fingerprints that emerge in their jailbreak responses. Empirical evaluations across multiple aligned LLMs reveal that persuasion-aware prompts significantly bypass safeguards, demonstrating their potential to induce jailbreak behaviors. This work underscores the importance of cross-disciplinary insight in addressing the evolving challenges of LLM safety. The code and data are available.
format Preprint
id arxiv_https___arxiv_org_abs_2510_21983
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
Noughabi, Havva Alizadeh
Serbanescu, Julien
Zarrinkalam, Fattane
Dehghantanha, Ali
Computation and Language
Artificial Intelligence
Despite recent advances, Large Language Models remain vulnerable to jailbreak attacks that bypass alignment safeguards and elicit harmful outputs. While prior research has proposed various attack strategies differing in human readability and transferability, little attention has been paid to the linguistic and psychological mechanisms that may influence a model's susceptibility to such attacks. In this paper, we examine an interdisciplinary line of research that leverages foundational theories of persuasion from the social sciences to craft adversarial prompts capable of circumventing alignment constraints in LLMs. Drawing on well-established persuasive strategies, we hypothesize that LLMs, having been trained on large-scale human-generated text, may respond more compliantly to prompts with persuasive structures. Furthermore, we investigate whether LLMs themselves exhibit distinct persuasive fingerprints that emerge in their jailbreak responses. Empirical evaluations across multiple aligned LLMs reveal that persuasion-aware prompts significantly bypass safeguards, demonstrating their potential to induce jailbreak behaviors. This work underscores the importance of cross-disciplinary insight in addressing the evolving challenges of LLM safety. The code and data are available.
title Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.21983