Adversarial Robustness for Visual Grounding of Multimodal Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gao, Kuofeng, Bai, Yang, Bai, Jiawang, Yang, Yong, Xia, Shu-Tao
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914798700068864
author Gao, Kuofeng
Bai, Yang
Bai, Jiawang
Yang, Yong
Xia, Shu-Tao
author_facet Gao, Kuofeng
Bai, Yang
Bai, Jiawang
Yang, Yong
Xia, Shu-Tao
contents Multi-modal Large Language Models (MLLMs) have recently achieved enhanced performance across various vision-language tasks including visual grounding capabilities. However, the adversarial robustness of visual grounding remains unexplored in MLLMs. To fill this gap, we use referring expression comprehension (REC) as an example task in visual grounding and propose three adversarial attack paradigms as follows. Firstly, untargeted adversarial attacks induce MLLMs to generate incorrect bounding boxes for each object. Besides, exclusive targeted adversarial attacks cause all generated outputs to the same target bounding box. In addition, permuted targeted adversarial attacks aim to permute all bounding boxes among different objects within a single image. Extensive experiments demonstrate that the proposed methods can successfully attack visual grounding capabilities of MLLMs. Our methods not only provide a new perspective for designing novel attacks but also serve as a strong baseline for improving the adversarial robustness for visual grounding of MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2405_09981
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
Gao, Kuofeng
Bai, Yang
Bai, Jiawang
Yang, Yong
Xia, Shu-Tao
Computer Vision and Pattern Recognition
Multi-modal Large Language Models (MLLMs) have recently achieved enhanced performance across various vision-language tasks including visual grounding capabilities. However, the adversarial robustness of visual grounding remains unexplored in MLLMs. To fill this gap, we use referring expression comprehension (REC) as an example task in visual grounding and propose three adversarial attack paradigms as follows. Firstly, untargeted adversarial attacks induce MLLMs to generate incorrect bounding boxes for each object. Besides, exclusive targeted adversarial attacks cause all generated outputs to the same target bounding box. In addition, permuted targeted adversarial attacks aim to permute all bounding boxes among different objects within a single image. Extensive experiments demonstrate that the proposed methods can successfully attack visual grounding capabilities of MLLMs. Our methods not only provide a new perspective for designing novel attacks but also serve as a strong baseline for improving the adversarial robustness for visual grounding of MLLMs.
title Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.09981