CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Guanghao, Qiu, Panjia, Fan, Mingyuan, Chen, Cen, Chu, Mingyuan, Zhang, Xin, Zhou, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910831215640576
author Zhou, Guanghao
Qiu, Panjia
Fan, Mingyuan
Chen, Cen
Chu, Mingyuan
Zhang, Xin
Zhou, Jun
author_facet Zhou, Guanghao
Qiu, Panjia
Fan, Mingyuan
Chen, Cen
Chu, Mingyuan
Zhang, Xin
Zhou, Jun
contents Despite explicit alignment efforts for large language models (LLMs), they can still be exploited to trigger unintended behaviors, a phenomenon known as "jailbreaking." Current jailbreak attack methods mainly focus on discrete prompt manipulations targeting closed-source LLMs, relying on manually crafted prompt templates and persuasion rules. However, as the capabilities of open-source LLMs improve, ensuring their safety becomes increasingly crucial. In such an environment, the accessibility of model parameters and gradient information by potential attackers exacerbates the severity of jailbreak threats. To address this research gap, we propose a novel \underline{C}ontext-\underline{C}oherent \underline{J}ailbreak \underline{A}ttack (CCJA). We define jailbreak attacks as an optimization problem within the embedding space of masked language models. Through combinatorial optimization, we effectively balance the jailbreak attack success rate with semantic coherence. Extensive evaluations show that our method not only maintains semantic consistency but also surpasses state-of-the-art baselines in attack effectiveness. Additionally, by integrating semantically coherent jailbreak prompts generated by our method into widely used black-box methodologies, we observe a notable enhancement in their success rates when targeting closed-source commercial LLMs. This highlights the security threat posed by open-source LLMs to commercial counterparts. We will open-source our code if the paper is accepted.
format Preprint
id arxiv_https___arxiv_org_abs_2502_11379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
Zhou, Guanghao
Qiu, Panjia
Fan, Mingyuan
Chen, Cen
Chu, Mingyuan
Zhang, Xin
Zhou, Jun
Cryptography and Security
Artificial Intelligence
Computation and Language
Despite explicit alignment efforts for large language models (LLMs), they can still be exploited to trigger unintended behaviors, a phenomenon known as "jailbreaking." Current jailbreak attack methods mainly focus on discrete prompt manipulations targeting closed-source LLMs, relying on manually crafted prompt templates and persuasion rules. However, as the capabilities of open-source LLMs improve, ensuring their safety becomes increasingly crucial. In such an environment, the accessibility of model parameters and gradient information by potential attackers exacerbates the severity of jailbreak threats. To address this research gap, we propose a novel \underline{C}ontext-\underline{C}oherent \underline{J}ailbreak \underline{A}ttack (CCJA). We define jailbreak attacks as an optimization problem within the embedding space of masked language models. Through combinatorial optimization, we effectively balance the jailbreak attack success rate with semantic coherence. Extensive evaluations show that our method not only maintains semantic consistency but also surpasses state-of-the-art baselines in attack effectiveness. Additionally, by integrating semantically coherent jailbreak prompts generated by our method into widely used black-box methodologies, we observe a notable enhancement in their success rates when targeting closed-source commercial LLMs. This highlights the security threat posed by open-source LLMs to commercial counterparts. We will open-source our code if the paper is accepted.
title CCJA: Context-Coherent Jailbreak Attack for Aligned Large Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2502.11379