HauntAttack: When Attack Follows Reasoning as a Shadow

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ma, Jingyuan, Li, Rui, Li, Zheng, Liu, Junfeng, Xia, Heming, Sha, Lei, Sui, Zhifang
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909864443248640
author Ma, Jingyuan
Li, Rui
Li, Zheng
Liu, Junfeng
Xia, Heming
Sha, Lei
Sui, Zhifang
author_facet Ma, Jingyuan
Li, Rui
Li, Zheng
Liu, Junfeng
Xia, Heming
Sha, Lei
Sui, Zhifang
contents Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However, the enhancement of reasoning abilities and the exposure of internal reasoning processes introduce new safety vulnerabilities. A critical question arises: when reasoning becomes intertwined with harmfulness, will LRMs become more vulnerable to jailbreaks in reasoning mode? To investigate this, we introduce HauntAttack, a novel and general-purpose black-box adversarial attack framework that systematically embeds harmful instructions into reasoning questions. Specifically, we modify key reasoning conditions in existing questions with harmful instructions, thereby constructing a reasoning pathway that guides the model step by step toward unsafe outputs. We evaluate HauntAttack on 11 LRMs and observe an average attack success rate of 70\%, achieving up to 12 percentage points of absolute improvement over the strongest prior baseline. Our further analysis reveals that even advanced safety-aligned models remain highly susceptible to reasoning-based attacks, offering insights into the urgent challenge of balancing reasoning capability and safety in future model development.
format Preprint
id arxiv_https___arxiv_org_abs_2506_07031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HauntAttack: When Attack Follows Reasoning as a Shadow
Ma, Jingyuan
Li, Rui
Li, Zheng
Liu, Junfeng
Xia, Heming
Sha, Lei
Sui, Zhifang
Cryptography and Security
Artificial Intelligence
Computation and Language
Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However, the enhancement of reasoning abilities and the exposure of internal reasoning processes introduce new safety vulnerabilities. A critical question arises: when reasoning becomes intertwined with harmfulness, will LRMs become more vulnerable to jailbreaks in reasoning mode? To investigate this, we introduce HauntAttack, a novel and general-purpose black-box adversarial attack framework that systematically embeds harmful instructions into reasoning questions. Specifically, we modify key reasoning conditions in existing questions with harmful instructions, thereby constructing a reasoning pathway that guides the model step by step toward unsafe outputs. We evaluate HauntAttack on 11 LRMs and observe an average attack success rate of 70\%, achieving up to 12 percentage points of absolute improvement over the strongest prior baseline. Our further analysis reveals that even advanced safety-aligned models remain highly susceptible to reasoning-based attacks, offering insights into the urgent challenge of balancing reasoning capability and safety in future model development.
title HauntAttack: When Attack Follows Reasoning as a Shadow
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.07031