ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cheng, Siyang, Liu, Gaotian, Mei, Rui, Wang, Yilin, Zhang, Kejia, Wei, Kaishuo, Yu, Yuqi, Wen, Weiping, Wu, Xiaojie, Liu, Junhua
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917085721919488
author Cheng, Siyang
Liu, Gaotian
Mei, Rui
Wang, Yilin
Zhang, Kejia
Wei, Kaishuo
Yu, Yuqi
Wen, Weiping
Wu, Xiaojie
Liu, Junhua
author_facet Cheng, Siyang
Liu, Gaotian
Mei, Rui
Wang, Yilin
Zhang, Kejia
Wei, Kaishuo
Yu, Yuqi
Wen, Weiping
Wu, Xiaojie
Liu, Junhua
contents The rapid adoption of large language models (LLMs) has brought both transformative applications and new security risks, including jailbreak attacks that bypass alignment safeguards to elicit harmful outputs. Existing automated jailbreak generation approaches e.g. AutoDAN, suffer from limited mutation diversity, shallow fitness evaluation, and fragile keyword-based detection. To address these limitations, we propose ForgeDAN, a novel evolutionary framework for generating semantically coherent and highly effective adversarial prompts against aligned LLMs. First, ForgeDAN introduces multi-strategy textual perturbations across \textit{character, word, and sentence-level} operations to enhance attack diversity; then we employ interpretable semantic fitness evaluation based on a text similarity model to guide the evolutionary process toward semantically relevant and harmful outputs; finally, ForgeDAN integrates dual-dimensional jailbreak judgment, leveraging an LLM-based classifier to jointly assess model compliance and output harmfulness, thereby reducing false positives and improving detection effectiveness. Our evaluation demonstrates ForgeDAN achieves high jailbreaking success rates while maintaining naturalness and stealth, outperforming existing SOTA solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2511_13548
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
Cheng, Siyang
Liu, Gaotian
Mei, Rui
Wang, Yilin
Zhang, Kejia
Wei, Kaishuo
Yu, Yuqi
Wen, Weiping
Wu, Xiaojie
Liu, Junhua
Cryptography and Security
Artificial Intelligence
Computation and Language
The rapid adoption of large language models (LLMs) has brought both transformative applications and new security risks, including jailbreak attacks that bypass alignment safeguards to elicit harmful outputs. Existing automated jailbreak generation approaches e.g. AutoDAN, suffer from limited mutation diversity, shallow fitness evaluation, and fragile keyword-based detection. To address these limitations, we propose ForgeDAN, a novel evolutionary framework for generating semantically coherent and highly effective adversarial prompts against aligned LLMs. First, ForgeDAN introduces multi-strategy textual perturbations across \textit{character, word, and sentence-level} operations to enhance attack diversity; then we employ interpretable semantic fitness evaluation based on a text similarity model to guide the evolutionary process toward semantically relevant and harmful outputs; finally, ForgeDAN integrates dual-dimensional jailbreak judgment, leveraging an LLM-based classifier to jointly assess model compliance and output harmfulness, thereby reducing false positives and improving detection effectiveness. Our evaluation demonstrates ForgeDAN achieves high jailbreaking success rates while maintaining naturalness and stealth, outperforming existing SOTA solutions.
title ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.13548