Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ji, Haoxuan, Lin, Zheng, Niu, Zhenxing, Gao, Xinbo, Hua, Gang
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908680263303168
author Ji, Haoxuan
Lin, Zheng
Niu, Zhenxing
Gao, Xinbo
Hua, Gang
author_facet Ji, Haoxuan
Lin, Zheng
Niu, Zhenxing
Gao, Xinbo
Hua, Gang
contents This paper focuses on jailbreaking attacks against large language models (LLMs), eliciting them to generate objectionable content in response to harmful user queries. Unlike previous LLM-jailbreak methods that directly orient to LLMs, our approach begins by constructing a multimodal large language model (MLLM) built upon the target LLM. Subsequently, we perform an efficient MLLM jailbreak and obtain a jailbreaking embedding. Finally, we convert the embedding into a textual jailbreaking suffix to carry out the jailbreak of target LLM. Compared to the direct LLM-jailbreak methods, our indirect jailbreaking approach is more efficient, as MLLMs are more vulnerable to jailbreak than pure LLM. Additionally, to improve the attack success rate of jailbreak, we propose an image-text semantic matching scheme to identify a suitable initial input. Extensive experiments demonstrate that our approach surpasses current state-of-the-art jailbreak methods in terms of both efficiency and effectiveness. Moreover, our approach exhibits superior cross-class generalization abilities.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20015
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
Ji, Haoxuan
Lin, Zheng
Niu, Zhenxing
Gao, Xinbo
Hua, Gang
Artificial Intelligence
Computation and Language
This paper focuses on jailbreaking attacks against large language models (LLMs), eliciting them to generate objectionable content in response to harmful user queries. Unlike previous LLM-jailbreak methods that directly orient to LLMs, our approach begins by constructing a multimodal large language model (MLLM) built upon the target LLM. Subsequently, we perform an efficient MLLM jailbreak and obtain a jailbreaking embedding. Finally, we convert the embedding into a textual jailbreaking suffix to carry out the jailbreak of target LLM. Compared to the direct LLM-jailbreak methods, our indirect jailbreaking approach is more efficient, as MLLMs are more vulnerable to jailbreak than pure LLM. Additionally, to improve the attack success rate of jailbreak, we propose an image-text semantic matching scheme to identify a suitable initial input. Extensive experiments demonstrate that our approach surpasses current state-of-the-art jailbreak methods in terms of both efficiency and effectiveness. Moreover, our approach exhibits superior cross-class generalization abilities.
title Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2405.20015