HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915567760310272 |
|---|---|
| author | Narula, Sidhant Asl, Javad Rafiei Ghasemigol, Mohammad Blanco, Eduardo Takabi, Daniel |
| author_facet | Narula, Sidhant Asl, Javad Rafiei Ghasemigol, Mohammad Blanco, Eduardo Takabi, Daniel |
| contents | Large Language Models (LLMs) remain vulnerable to multi-turn jailbreak attacks. We introduce HarmNet, a modular framework comprising ThoughtNet, a hierarchical semantic network; a feedback-driven Simulator for iterative query refinement; and a Network Traverser for real-time adaptive attack execution. HarmNet systematically explores and refines the adversarial space to uncover stealthy, high-success attack paths. Experiments across closed-source and open-source LLMs show that HarmNet outperforms state-of-the-art methods, achieving higher attack success rates. For example, on Mistral-7B, HarmNet achieves a 99.4% attack success rate, 13.9% higher than the best baseline. Index terms: jailbreak attacks; large language models; adversarial framework; query refinement. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_18728 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models Narula, Sidhant Asl, Javad Rafiei Ghasemigol, Mohammad Blanco, Eduardo Takabi, Daniel Cryptography and Security Artificial Intelligence Large Language Models (LLMs) remain vulnerable to multi-turn jailbreak attacks. We introduce HarmNet, a modular framework comprising ThoughtNet, a hierarchical semantic network; a feedback-driven Simulator for iterative query refinement; and a Network Traverser for real-time adaptive attack execution. HarmNet systematically explores and refines the adversarial space to uncover stealthy, high-success attack paths. Experiments across closed-source and open-source LLMs show that HarmNet outperforms state-of-the-art methods, achieving higher attack success rates. For example, on Mistral-7B, HarmNet achieves a 99.4% attack success rate, 13.9% higher than the best baseline. Index terms: jailbreak attacks; large language models; adversarial framework; query refinement. |
| title | HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models |
| topic | Cryptography and Security Artificial Intelligence |
| url | https://arxiv.org/abs/2510.18728 |