Defending LVLMs Against Vision Attacks through Partial-Perception Supervision

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Qi, Li, Tianlin, Guo, Qing, Wang, Dongxia, Lin, Yun, Liu, Yang, Dong, Jin Song
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908518710247424
author Zhou, Qi
Li, Tianlin
Guo, Qing
Wang, Dongxia
Lin, Yun
Liu, Yang
Dong, Jin Song
author_facet Zhou, Qi
Li, Tianlin
Guo, Qing
Wang, Dongxia
Lin, Yun
Liu, Yang
Dong, Jin Song
contents Recent studies have raised significant concerns regarding the vulnerability of Large Vision Language Models (LVLMs) to maliciously injected or perturbed input images, which can mislead their responses. Existing defense methods show that such vision attacks are sensitive to image modifications especially cropping, using majority voting across responses of modified images as corrected responses. However, these modifications often result in partial images and distort the semantics, which reduces response quality on clean images after voting. Instead of directly using responses from partial images for voting, we investigate using them to supervise the LVLM's responses to the original images. We propose a black-box, training-free method called DPS (Defense through Partial-Perception Supervision). In this approach, the model is prompted using the responses generated by a model that perceives only a partial image. With DPS, the model can adjust its response based on partial image understanding when under attack, while confidently maintaining its original response for clean input. Our findings show that the weak model can supervise the strong model: when faced with an attacked input, the strong model becomes less confident and adjusts its response based on the weak model's partial understanding, effectively defending against the attack. With clean input, it confidently maintains its original response. Empirical experiments show our method outperforms the baseline, cutting the average attack success rate by 76.3% across six datasets on three popular models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12722
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
Zhou, Qi
Li, Tianlin
Guo, Qing
Wang, Dongxia
Lin, Yun
Liu, Yang
Dong, Jin Song
Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
Recent studies have raised significant concerns regarding the vulnerability of Large Vision Language Models (LVLMs) to maliciously injected or perturbed input images, which can mislead their responses. Existing defense methods show that such vision attacks are sensitive to image modifications especially cropping, using majority voting across responses of modified images as corrected responses. However, these modifications often result in partial images and distort the semantics, which reduces response quality on clean images after voting. Instead of directly using responses from partial images for voting, we investigate using them to supervise the LVLM's responses to the original images. We propose a black-box, training-free method called DPS (Defense through Partial-Perception Supervision). In this approach, the model is prompted using the responses generated by a model that perceives only a partial image. With DPS, the model can adjust its response based on partial image understanding when under attack, while confidently maintaining its original response for clean input. Our findings show that the weak model can supervise the strong model: when faced with an attacked input, the strong model becomes less confident and adjusts its response based on the weak model's partial understanding, effectively defending against the attack. With clean input, it confidently maintains its original response. Empirical experiments show our method outperforms the baseline, cutting the average attack success rate by 76.3% across six datasets on three popular models.
title Defending LVLMs Against Vision Attacks through Partial-Perception Supervision
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2412.12722