Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Huang, Yao, Sun, Yitong, Ruan, Shouwei, Zhang, Yichi, Dong, Yinpeng, Wei, Xingxing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912399934619648
author Huang, Yao
Sun, Yitong
Ruan, Shouwei
Zhang, Yichi
Dong, Yinpeng
Wei, Xingxing
author_facet Huang, Yao
Sun, Yitong
Ruan, Shouwei
Zhang, Yichi
Dong, Yinpeng
Wei, Xingxing
contents Large Language Models (LLMs), despite advanced general capabilities, still suffer from numerous safety risks, especially jailbreak attacks that bypass safety protocols. Understanding these vulnerabilities through black-box jailbreak attacks, which better reflect real-world scenarios, offers critical insights into model robustness. While existing methods have shown improvements through various prompt engineering techniques, their success remains limited against safety-aligned models, overlooking a more fundamental problem: the effectiveness is inherently bounded by the predefined strategy spaces. However, expanding this space presents significant challenges in both systematically capturing essential attack patterns and efficiently navigating the increased complexity. To better explore the potential of expanding the strategy space, we address these challenges through a novel framework that decomposes jailbreak strategies into essential components based on the Elaboration Likelihood Model (ELM) theory and develops genetic-based optimization with intention evaluation mechanisms. To be striking, our experiments reveal unprecedented jailbreak capabilities by expanding the strategy space: we achieve over 90% success rate on Claude-3.5 where prior methods completely fail, while demonstrating strong cross-model transferability and surpassing specialized safeguard models in evaluation accuracy. The code is open-sourced at: https://github.com/Aries-iai/CL-GSO.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21277
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
Huang, Yao
Sun, Yitong
Ruan, Shouwei
Zhang, Yichi
Dong, Yinpeng
Wei, Xingxing
Cryptography and Security
Artificial Intelligence
Computation and Language
Large Language Models (LLMs), despite advanced general capabilities, still suffer from numerous safety risks, especially jailbreak attacks that bypass safety protocols. Understanding these vulnerabilities through black-box jailbreak attacks, which better reflect real-world scenarios, offers critical insights into model robustness. While existing methods have shown improvements through various prompt engineering techniques, their success remains limited against safety-aligned models, overlooking a more fundamental problem: the effectiveness is inherently bounded by the predefined strategy spaces. However, expanding this space presents significant challenges in both systematically capturing essential attack patterns and efficiently navigating the increased complexity. To better explore the potential of expanding the strategy space, we address these challenges through a novel framework that decomposes jailbreak strategies into essential components based on the Elaboration Likelihood Model (ELM) theory and develops genetic-based optimization with intention evaluation mechanisms. To be striking, our experiments reveal unprecedented jailbreak capabilities by expanding the strategy space: we achieve over 90% success rate on Claude-3.5 where prior methods completely fail, while demonstrating strong cross-model transferability and surpassing specialized safeguard models in evaluation accuracy. The code is open-sourced at: https://github.com/Aries-iai/CL-GSO.
title Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2505.21277