Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Zihao, Liu, Yi, Deng, Gelei, Wang, Kailong, Li, Yuekang, Shi, Ling, Picek, Stjepan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910534961463296
author Xu, Zihao
Liu, Yi
Deng, Gelei
Wang, Kailong
Li, Yuekang
Shi, Ling
Picek, Stjepan
author_facet Xu, Zihao
Liu, Yi
Deng, Gelei
Wang, Kailong
Li, Yuekang
Shi, Ling
Picek, Stjepan
contents Security concerns for large language models (LLMs) have recently escalated, focusing on thwarting jailbreaking attempts in discrete prompts. However, the exploration of jailbreak vulnerabilities arising from continuous embeddings has been limited, as prior approaches primarily involved appending discrete or continuous suffixes to inputs. Our study presents a novel channel for conducting direct attacks on LLM inputs, eliminating the need for suffix addition or specific questions provided that the desired output is predefined. We additionally observe that extensive iterations often lead to overfitting, characterized by repetition in the output. To counteract this, we propose a simple yet effective strategy named CLIP. Our experiments show that for an input length of 40 at iteration 1000, applying CLIP improves the ASR from 62% to 83%
format Preprint
id arxiv_https___arxiv_org_abs_2407_13796
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
Xu, Zihao
Liu, Yi
Deng, Gelei
Wang, Kailong
Li, Yuekang
Shi, Ling
Picek, Stjepan
Cryptography and Security
Artificial Intelligence
Computation and Language
Security concerns for large language models (LLMs) have recently escalated, focusing on thwarting jailbreaking attempts in discrete prompts. However, the exploration of jailbreak vulnerabilities arising from continuous embeddings has been limited, as prior approaches primarily involved appending discrete or continuous suffixes to inputs. Our study presents a novel channel for conducting direct attacks on LLM inputs, eliminating the need for suffix addition or specific questions provided that the desired output is predefined. We additionally observe that extensive iterations often lead to overfitting, characterized by repetition in the output. To counteract this, we propose a simple yet effective strategy named CLIP. Our experiments show that for an input length of 40 at iteration 1000, applying CLIP improves the ASR from 62% to 83%
title Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2407.13796