Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Yunhao, Wang, Xin, Li, Juncheng, Wang, Yixu, Li, Jie, Teng, Yan, Wang, Yingchun, Ma, Xingjun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917507083796480
author Chen, Yunhao
Wang, Xin
Li, Juncheng
Wang, Yixu
Li, Jie
Teng, Yan
Wang, Yingchun
Ma, Xingjun
author_facet Chen, Yunhao
Wang, Xin
Li, Juncheng
Wang, Yixu
Li, Jie
Teng, Yan
Wang, Yingchun
Ma, Xingjun
contents Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophisticated, yet many still formulate attack optimization primarily in the prompt space. In other words, these methods mainly search for better attack wording or better strategy choices, but they do not search over executable code. By moving the search into code space, we can optimize not only the final attack prompt, but also the procedure that generates it, including execution flow, reusable logic, branching, and failure-driven repair. To overcome this gap, we introduce EvoSynth, an autonomous multi-agent framework that shifts the optimization space from prompts to executable code. Instead of refining prompts directly, EvoSynth employs a multi-agent system to autonomously engineer, evolve, and execute code-based attack algorithms. Crucially, it features a code-level self-correction loop, allowing it to iteratively rewrite the code-based algorithm in response to target-model feedback and failed attempts. Through extensive experiments, we demonstrate that EvoSynth achieves an 85.5\% Attack Success Rate (ASR) against highly robust models like Claude-Sonnet-4.5 and a 95.9\% average ASR across evaluated targets, while generating attacks that are significantly more diverse than those from existing methods. We release our framework to facilitate future research on evolutionary synthesis in executable code space.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12710
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
Chen, Yunhao
Wang, Xin
Li, Juncheng
Wang, Yixu
Li, Jie
Teng, Yan
Wang, Yingchun
Ma, Xingjun
Computation and Language
Cryptography and Security
Automated red teaming frameworks for Large Language Models (LLMs) have become increasingly sophisticated, yet many still formulate attack optimization primarily in the prompt space. In other words, these methods mainly search for better attack wording or better strategy choices, but they do not search over executable code. By moving the search into code space, we can optimize not only the final attack prompt, but also the procedure that generates it, including execution flow, reusable logic, branching, and failure-driven repair. To overcome this gap, we introduce EvoSynth, an autonomous multi-agent framework that shifts the optimization space from prompts to executable code. Instead of refining prompts directly, EvoSynth employs a multi-agent system to autonomously engineer, evolve, and execute code-based attack algorithms. Crucially, it features a code-level self-correction loop, allowing it to iteratively rewrite the code-based algorithm in response to target-model feedback and failed attempts. Through extensive experiments, we demonstrate that EvoSynth achieves an 85.5\% Attack Success Rate (ASR) against highly robust models like Claude-Sonnet-4.5 and a 95.9\% average ASR across evaluated targets, while generating attacks that are significantly more diverse than those from existing methods. We release our framework to facilitate future research on evolutionary synthesis in executable code space.
title Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
topic Computation and Language
Cryptography and Security
url https://arxiv.org/abs/2511.12710