FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gong, Yichen, Ran, Delong, Liu, Jinyuan, Wang, Conglei, Cong, Tianshuo, Wang, Anyu, Duan, Sisi, Wang, Xiaoyun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915108344561664
author Gong, Yichen
Ran, Delong
Liu, Jinyuan
Wang, Conglei
Cong, Tianshuo
Wang, Anyu
Duan, Sisi
Wang, Xiaoyun
author_facet Gong, Yichen
Ran, Delong
Liu, Jinyuan
Wang, Conglei
Cong, Tianshuo
Wang, Anyu
Duan, Sisi
Wang, Xiaoyun
contents Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language Models (LLMs) by assimilating additional modalities (e.g., images). Despite this advancement, the safety of LVLMs remains adequately underexplored, with a potential overreliance on the safety assurances purported by their underlying LLMs. In this paper, we propose FigStep, a straightforward yet effective black-box jailbreak algorithm against LVLMs. Instead of feeding textual harmful instructions directly, FigStep converts the prohibited content into images through typography to bypass the safety alignment. The experimental results indicate that FigStep can achieve an average attack success rate of 82.50% on six promising open-source LVLMs. Not merely to demonstrate the efficacy of FigStep, we conduct comprehensive ablation studies and analyze the distribution of the semantic embeddings to uncover that the reason behind the success of FigStep is the deficiency of safety alignment for visual embeddings. Moreover, we compare FigStep with five text-only jailbreaks and four image-based jailbreaks to demonstrate the superiority of FigStep, i.e., negligible attack costs and better attack performance. Above all, our work reveals that current LVLMs are vulnerable to jailbreak attacks, which highlights the necessity of novel cross-modality safety alignment techniques. Our code and datasets are available at https://github.com/ThuCCSLab/FigStep .
format Preprint
id arxiv_https___arxiv_org_abs_2311_05608
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
Gong, Yichen
Ran, Delong
Liu, Jinyuan
Wang, Conglei
Cong, Tianshuo
Wang, Anyu
Duan, Sisi
Wang, Xiaoyun
Cryptography and Security
Artificial Intelligence
Computation and Language
Large Vision-Language Models (LVLMs) signify a groundbreaking paradigm shift within the Artificial Intelligence (AI) community, extending beyond the capabilities of Large Language Models (LLMs) by assimilating additional modalities (e.g., images). Despite this advancement, the safety of LVLMs remains adequately underexplored, with a potential overreliance on the safety assurances purported by their underlying LLMs. In this paper, we propose FigStep, a straightforward yet effective black-box jailbreak algorithm against LVLMs. Instead of feeding textual harmful instructions directly, FigStep converts the prohibited content into images through typography to bypass the safety alignment. The experimental results indicate that FigStep can achieve an average attack success rate of 82.50% on six promising open-source LVLMs. Not merely to demonstrate the efficacy of FigStep, we conduct comprehensive ablation studies and analyze the distribution of the semantic embeddings to uncover that the reason behind the success of FigStep is the deficiency of safety alignment for visual embeddings. Moreover, we compare FigStep with five text-only jailbreaks and four image-based jailbreaks to demonstrate the superiority of FigStep, i.e., negligible attack costs and better attack performance. Above all, our work reveals that current LVLMs are vulnerable to jailbreak attacks, which highlights the necessity of novel cross-modality safety alignment techniques. Our code and datasets are available at https://github.com/ThuCCSLab/FigStep .
title FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2311.05608