Enhancing Jailbreak Attacks on LLMs via Persona Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zheng, Zhao, Peilin, Ye, Deheng, Wang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918407881883648
author Zhang, Zheng
Zhao, Peilin
Ye, Deheng
Wang, Hao
author_facet Zhang, Zheng
Zhao, Peilin
Ye, Deheng
Wang, Hao
contents Jailbreak attacks aim to exploit large language models (LLMs) by inducing them to generate harmful content, thereby revealing their vulnerabilities. Understanding and addressing these attacks is crucial for advancing the field of LLM safety. Previous jailbreak approaches have mainly focused on direct manipulations of harmful intent, with limited attention to the impact of persona prompts. In this study, we systematically explore the efficacy of persona prompts in compromising LLM defenses. We propose a genetic algorithm-based method that automatically crafts persona prompts to bypass LLM's safety mechanisms. Our experiments reveal that: (1) our evolved persona prompts reduce refusal rates by 50-70% across multiple LLMs, and (2) these prompts demonstrate synergistic effects when combined with existing attack methods, increasing success rates by 10-20%. Our code and data are available at https://github.com/CjangCjengh/Generic_Persona.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22171
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Jailbreak Attacks on LLMs via Persona Prompts
Zhang, Zheng
Zhao, Peilin
Ye, Deheng
Wang, Hao
Cryptography and Security
Artificial Intelligence
Jailbreak attacks aim to exploit large language models (LLMs) by inducing them to generate harmful content, thereby revealing their vulnerabilities. Understanding and addressing these attacks is crucial for advancing the field of LLM safety. Previous jailbreak approaches have mainly focused on direct manipulations of harmful intent, with limited attention to the impact of persona prompts. In this study, we systematically explore the efficacy of persona prompts in compromising LLM defenses. We propose a genetic algorithm-based method that automatically crafts persona prompts to bypass LLM's safety mechanisms. Our experiments reveal that: (1) our evolved persona prompts reduce refusal rates by 50-70% across multiple LLMs, and (2) these prompts demonstrate synergistic effects when combined with existing attack methods, increasing success rates by 10-20%. Our code and data are available at https://github.com/CjangCjengh/Generic_Persona.
title Enhancing Jailbreak Attacks on LLMs via Persona Prompts
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2507.22171