Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Mao, Yanxu, Liu, Peipei, Cui, Tiehan, Yan, Zhaoteng, Liu, Congying, You, Datao
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910973506355200
author Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Yan, Zhaoteng
Liu, Congying
You, Datao
author_facet Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Yan, Zhaoteng
Liu, Congying
You, Datao
contents Large language models (LLMs) are widely applied in various fields of society due to their powerful reasoning, understanding, and generation capabilities. However, the security issues associated with these models are becoming increasingly severe. Jailbreaking attacks, as an important method for detecting vulnerabilities in LLMs, have been explored by researchers who attempt to induce these models to generate harmful content through various attack methods. Nevertheless, existing jailbreaking methods face numerous limitations, such as excessive query counts, limited coverage of jailbreak modalities, low attack success rates, and simplistic evaluation methods. To overcome these constraints, this paper proposes a multimodal jailbreaking method: JMLLM. This method integrates multiple strategies to perform comprehensive jailbreak attacks across text, visual, and auditory modalities. Additionally, we contribute a new and comprehensive dataset for multimodal jailbreaking research: TriJail, which includes jailbreak prompts for all three modalities. Experiments on the TriJail dataset and the benchmark dataset AdvBench, conducted on 13 popular LLMs, demonstrate advanced attack success rates and significant reduction in time overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16555
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models
Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Yan, Zhaoteng
Liu, Congying
You, Datao
Computation and Language
Large language models (LLMs) are widely applied in various fields of society due to their powerful reasoning, understanding, and generation capabilities. However, the security issues associated with these models are becoming increasingly severe. Jailbreaking attacks, as an important method for detecting vulnerabilities in LLMs, have been explored by researchers who attempt to induce these models to generate harmful content through various attack methods. Nevertheless, existing jailbreaking methods face numerous limitations, such as excessive query counts, limited coverage of jailbreak modalities, low attack success rates, and simplistic evaluation methods. To overcome these constraints, this paper proposes a multimodal jailbreaking method: JMLLM. This method integrates multiple strategies to perform comprehensive jailbreak attacks across text, visual, and auditory modalities. Additionally, we contribute a new and comprehensive dataset for multimodal jailbreaking research: TriJail, which includes jailbreak prompts for all three modalities. Experiments on the TriJail dataset and the benchmark dataset AdvBench, conducted on 13 popular LLMs, demonstrate advanced attack success rates and significant reduction in time overhead.
title Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2412.16555