Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Yuting, Liu, Chengyuan, Feng, Yifeng, Wu, Yiquan, Wu, Chao, Wu, Fei, Kuang, Kun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915312883990528
author Huang, Yuting
Liu, Chengyuan
Feng, Yifeng
Wu, Yiquan
Wu, Chao
Wu, Fei
Kuang, Kun
author_facet Huang, Yuting
Liu, Chengyuan
Feng, Yifeng
Wu, Yiquan
Wu, Chao
Wu, Fei
Kuang, Kun
contents As Large Language Models (LLMs) are widely applied in various domains, the safety of LLMs is increasingly attracting attention to avoid their powerful capabilities being misused. Existing jailbreak methods create a forced instruction-following scenario, or search adversarial prompts with prefix or suffix tokens to achieve a specific representation manually or automatically. However, they suffer from low efficiency and explicit jailbreak patterns, far from the real deployment of mass attacks to LLMs. In this paper, we point out that simply rewriting the original instruction can achieve a jailbreak, and we find that this rewriting approach is learnable and transferable. We propose the Rewrite to Jailbreak (R2J) approach, a transferable black-box jailbreak method to attack LLMs by iteratively exploring the weakness of the LLMs and automatically improving the attacking strategy. The jailbreak is more efficient and hard to identify since no additional features are introduced. Extensive experiments and analysis demonstrate the effectiveness of R2J, and we find that the jailbreak is also transferable to multiple datasets and various types of models with only a few queries. We hope our work motivates further investigation of LLM safety. The code can be found at https://github.com/ythuang02/R2J/.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11084
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction
Huang, Yuting
Liu, Chengyuan
Feng, Yifeng
Wu, Yiquan
Wu, Chao
Wu, Fei
Kuang, Kun
Computation and Language
As Large Language Models (LLMs) are widely applied in various domains, the safety of LLMs is increasingly attracting attention to avoid their powerful capabilities being misused. Existing jailbreak methods create a forced instruction-following scenario, or search adversarial prompts with prefix or suffix tokens to achieve a specific representation manually or automatically. However, they suffer from low efficiency and explicit jailbreak patterns, far from the real deployment of mass attacks to LLMs. In this paper, we point out that simply rewriting the original instruction can achieve a jailbreak, and we find that this rewriting approach is learnable and transferable. We propose the Rewrite to Jailbreak (R2J) approach, a transferable black-box jailbreak method to attack LLMs by iteratively exploring the weakness of the LLMs and automatically improving the attacking strategy. The jailbreak is more efficient and hard to identify since no additional features are introduced. Extensive experiments and analysis demonstrate the effectiveness of R2J, and we find that the jailbreak is also transferable to multiple datasets and various types of models with only a few queries. We hope our work motivates further investigation of LLM safety. The code can be found at https://github.com/ythuang02/R2J/.
title Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction
topic Computation and Language
url https://arxiv.org/abs/2502.11084