Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hao, Shuyang, Wang, Yiwei, Hooi, Bryan, Yang, Ming-Hsuan, Liu, Jun, Tang, Chengcheng, Huang, Zi, Cai, Yujun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915198613323776
author Hao, Shuyang
Wang, Yiwei
Hooi, Bryan
Yang, Ming-Hsuan
Liu, Jun
Tang, Chengcheng
Huang, Zi
Cai, Yujun
author_facet Hao, Shuyang
Wang, Yiwei
Hooi, Bryan
Yang, Ming-Hsuan
Liu, Jun
Tang, Chengcheng
Huang, Zi
Cai, Yujun
contents Deploying large vision-language models (LVLMs) introduces a unique vulnerability: susceptibility to malicious attacks via visual inputs. However, existing defense methods suffer from two key limitations: (1) They solely focus on textual defenses, fail to directly address threats in the visual domain where attacks originate, and (2) the additional processing steps often incur significant computational overhead or compromise model performance on benign tasks. Building on these insights, we propose ESIII (Embedding Security Instructions Into Images), a novel methodology for transforming the visual space from a source of vulnerability into an active defense mechanism. Initially, we embed security instructions into defensive images through gradient-based optimization, obtaining security instructions in the visual dimension. Subsequently, we integrate security instructions from visual and textual dimensions with the input query. The collaboration between security instructions from different dimensions ensures comprehensive security protection. Extensive experiments demonstrate that our approach effectively fortifies the robustness of LVLMs against such attacks while preserving their performance on standard benign tasks and incurring an imperceptible increase in time costs.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
Hao, Shuyang
Wang, Yiwei
Hooi, Bryan
Yang, Ming-Hsuan
Liu, Jun
Tang, Chengcheng
Huang, Zi
Cai, Yujun
Cryptography and Security
Deploying large vision-language models (LVLMs) introduces a unique vulnerability: susceptibility to malicious attacks via visual inputs. However, existing defense methods suffer from two key limitations: (1) They solely focus on textual defenses, fail to directly address threats in the visual domain where attacks originate, and (2) the additional processing steps often incur significant computational overhead or compromise model performance on benign tasks. Building on these insights, we propose ESIII (Embedding Security Instructions Into Images), a novel methodology for transforming the visual space from a source of vulnerability into an active defense mechanism. Initially, we embed security instructions into defensive images through gradient-based optimization, obtaining security instructions in the visual dimension. Subsequently, we integrate security instructions from visual and textual dimensions with the input query. The collaboration between security instructions from different dimensions ensures comprehensive security protection. Extensive experiments demonstrate that our approach effectively fortifies the robustness of LVLMs against such attacks while preserving their performance on standard benign tasks and incurring an imperceptible increase in time costs.
title Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
topic Cryptography and Security
url https://arxiv.org/abs/2503.11619