Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ding, Peng, Kuang, Jun, Sun, Wen, Wang, Zongyu, Cao, Xuezhi, Cai, Xunliang, Chen, Jiajun, Huang, Shujian
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918181648465920
author Ding, Peng
Kuang, Jun
Sun, Wen
Wang, Zongyu
Cao, Xuezhi
Cai, Xunliang
Chen, Jiajun
Huang, Shujian
author_facet Ding, Peng
Kuang, Jun
Sun, Wen
Wang, Zongyu
Cao, Xuezhi
Cai, Xunliang
Chen, Jiajun
Huang, Shujian
contents Large language models (LLMs) remain vulnerable to jailbreaking attacks despite their impressive capabilities. Investigating these weaknesses is crucial for robust safety mechanisms. Existing attacks primarily distract LLMs by introducing additional context or adversarial tokens, leaving the core harmful intent unchanged. In this paper, we introduce ISA (Intent Shift Attack), which obfuscates LLMs about the intent of the attacks. More specifically, we establish a taxonomy of intent transformations and leverage them to generate attacks that may be misperceived by LLMs as benign requests for information. Unlike prior methods relying on complex tokens or lengthy context, our approach only needs minimal edits to the original request, and yields natural, human-readable, and seemingly harmless prompts. Extensive experiments on both open-source and commercial LLMs show that ISA achieves over 70% improvement in attack success rate compared to direct harmful prompts. More critically, fine-tuning models on only benign data reformulated with ISA templates elevates success rates to nearly 100%. For defense, we evaluate existing methods and demonstrate their inadequacy against ISA, while exploring both training-free and training-based mitigation strategies. Our findings reveal fundamental challenges in intent inference for LLMs safety and underscore the need for more effective defenses. Our code and datasets are available at https://github.com/NJUNLP/ISA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00556
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
Ding, Peng
Kuang, Jun
Sun, Wen
Wang, Zongyu
Cao, Xuezhi
Cai, Xunliang
Chen, Jiajun
Huang, Shujian
Computation and Language
Large language models (LLMs) remain vulnerable to jailbreaking attacks despite their impressive capabilities. Investigating these weaknesses is crucial for robust safety mechanisms. Existing attacks primarily distract LLMs by introducing additional context or adversarial tokens, leaving the core harmful intent unchanged. In this paper, we introduce ISA (Intent Shift Attack), which obfuscates LLMs about the intent of the attacks. More specifically, we establish a taxonomy of intent transformations and leverage them to generate attacks that may be misperceived by LLMs as benign requests for information. Unlike prior methods relying on complex tokens or lengthy context, our approach only needs minimal edits to the original request, and yields natural, human-readable, and seemingly harmless prompts. Extensive experiments on both open-source and commercial LLMs show that ISA achieves over 70% improvement in attack success rate compared to direct harmful prompts. More critically, fine-tuning models on only benign data reformulated with ISA templates elevates success rates to nearly 100%. For defense, we evaluate existing methods and demonstrate their inadequacy against ISA, while exploring both training-free and training-based mitigation strategies. Our findings reveal fundamental challenges in intent inference for LLMs safety and underscore the need for more effective defenses. Our code and datasets are available at https://github.com/NJUNLP/ISA.
title Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
topic Computation and Language
url https://arxiv.org/abs/2511.00556