Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ahn, Jaewoo, Yun, Heeseung, Ko, Dayoon, Kim, Gunhee
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918037635989504
author Ahn, Jaewoo
Yun, Heeseung
Ko, Dayoon
Kim, Gunhee
author_facet Ahn, Jaewoo
Yun, Heeseung
Ko, Dayoon
Kim, Gunhee
contents While pre-trained multimodal representations (e.g., CLIP) have shown impressive capabilities, they exhibit significant compositional vulnerabilities leading to counterintuitive judgments. We introduce Multimodal Adversarial Compositionality (MAC), a benchmark that leverages large language models (LLMs) to generate deceptive text samples to exploit these vulnerabilities across different modalities and evaluates them through both sample-wise attack success rate and group-wise entropy-based diversity. To improve zero-shot methods, we propose a self-training approach that leverages rejection-sampling fine-tuning with diversity-promoting filtering, which enhances both attack success rate and sample diversity. Using smaller language models like Llama-3.1-8B, our approach demonstrates superior performance in revealing compositional vulnerabilities across various multimodal representations, including images, videos, and audios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22943
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
Ahn, Jaewoo
Yun, Heeseung
Ko, Dayoon
Kim, Gunhee
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
While pre-trained multimodal representations (e.g., CLIP) have shown impressive capabilities, they exhibit significant compositional vulnerabilities leading to counterintuitive judgments. We introduce Multimodal Adversarial Compositionality (MAC), a benchmark that leverages large language models (LLMs) to generate deceptive text samples to exploit these vulnerabilities across different modalities and evaluates them through both sample-wise attack success rate and group-wise entropy-based diversity. To improve zero-shot methods, we propose a self-training approach that leverages rejection-sampling fine-tuning with diversity-promoting filtering, which enhances both attack success rate and sample diversity. Using smaller language models like Llama-3.1-8B, our approach demonstrates superior performance in revealing compositional vulnerabilities across various multimodal representations, including images, videos, and audios.
title Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
url https://arxiv.org/abs/2505.22943