Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yuqi, Lu, Lin, Sun, Hanchi, Zhou, Pan, Sun, Lichao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911951964078080
author Zhou, Yuqi
Lu, Lin
Sun, Hanchi
Zhou, Pan
Sun, Lichao
author_facet Zhou, Yuqi
Lu, Lin
Sun, Hanchi
Zhou, Pan
Sun, Lichao
contents Jailbreak attacks on large language models (LLMs) involve inducing these models to generate harmful content that violates ethics or laws, posing a significant threat to LLM security. Current jailbreak attacks face two main challenges: low success rates due to defensive measures and high resource requirements for crafting specific prompts. This paper introduces Virtual Context, which leverages special tokens, previously overlooked in LLM security, to improve jailbreak attacks. Virtual Context addresses these challenges by significantly increasing the success rates of existing jailbreak methods and requiring minimal background knowledge about the target model, thus enhancing effectiveness in black-box settings without additional overhead. Comprehensive evaluations show that Virtual Context-assisted jailbreak attacks can improve the success rates of four widely used jailbreak methods by approximately 40% across various LLMs. Additionally, applying Virtual Context to original malicious behaviors still achieves a notable jailbreak effect. In summary, our research highlights the potential of special tokens in jailbreak attacks and recommends including this threat in red-teaming testing to comprehensively enhance LLM security.
format Preprint
id arxiv_https___arxiv_org_abs_2406_19845
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
Zhou, Yuqi
Lu, Lin
Sun, Hanchi
Zhou, Pan
Sun, Lichao
Cryptography and Security
Jailbreak attacks on large language models (LLMs) involve inducing these models to generate harmful content that violates ethics or laws, posing a significant threat to LLM security. Current jailbreak attacks face two main challenges: low success rates due to defensive measures and high resource requirements for crafting specific prompts. This paper introduces Virtual Context, which leverages special tokens, previously overlooked in LLM security, to improve jailbreak attacks. Virtual Context addresses these challenges by significantly increasing the success rates of existing jailbreak methods and requiring minimal background knowledge about the target model, thus enhancing effectiveness in black-box settings without additional overhead. Comprehensive evaluations show that Virtual Context-assisted jailbreak attacks can improve the success rates of four widely used jailbreak methods by approximately 40% across various LLMs. Additionally, applying Virtual Context to original malicious behaviors still achieves a notable jailbreak effect. In summary, our research highlights the potential of special tokens in jailbreak attacks and recommends including this threat in red-teaming testing to comprehensively enhance LLM security.
title Virtual Context: Enhancing Jailbreak Attacks with Special Token Injection
topic Cryptography and Security
url https://arxiv.org/abs/2406.19845