Saved in:
Bibliographic Details
Main Authors: Akbarian, Fatemeh, Baninajjar, Anahita, Zhang, Yingyi, Balashankar, Ananth, Aminifar, Amir
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.21893
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910151300087808
author Akbarian, Fatemeh
Baninajjar, Anahita
Zhang, Yingyi
Balashankar, Ananth
Aminifar, Amir
author_facet Akbarian, Fatemeh
Baninajjar, Anahita
Zhang, Yingyi
Balashankar, Ananth
Aminifar, Amir
contents Multi-modal foundation models align images, text, and other modalities in a shared embedding space but remain vulnerable to adversarial illusions [35], where imperceptible perturbations disrupt cross-modal alignment and mislead downstream tasks. To counteract the effects of adversarial illusions, we propose a task-agnostic mitigation mechanism that purifies the attacker's perturbed input using generative models, e.g., Variational Autoencoders (VAEs), to restore natural alignment. To further enhance the defense mechanism, we adopt a generative sampling strategy combined with a consensus-based aggregation scheme over the outcomes of the generated samples. Our experiments on ImageBind, a state-of-the-art multi-modal encoder, show that our approach substantially reduces the illusion attack success rates to near-zero and improves cross-modal alignment in unperturbed and perturbed input settings, providing an effective and task-agnostic defense against adversarial illusions. The code is available at https://github.com/fatemehakb/adversarial-illusions-mitigation.
format Preprint
id arxiv_https___arxiv_org_abs_2511_21893
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
Akbarian, Fatemeh
Baninajjar, Anahita
Zhang, Yingyi
Balashankar, Ananth
Aminifar, Amir
Machine Learning
Multi-modal foundation models align images, text, and other modalities in a shared embedding space but remain vulnerable to adversarial illusions [35], where imperceptible perturbations disrupt cross-modal alignment and mislead downstream tasks. To counteract the effects of adversarial illusions, we propose a task-agnostic mitigation mechanism that purifies the attacker's perturbed input using generative models, e.g., Variational Autoencoders (VAEs), to restore natural alignment. To further enhance the defense mechanism, we adopt a generative sampling strategy combined with a consensus-based aggregation scheme over the outcomes of the generated samples. Our experiments on ImageBind, a state-of-the-art multi-modal encoder, show that our approach substantially reduces the illusion attack success rates to near-zero and improves cross-modal alignment in unperturbed and perturbed input settings, providing an effective and task-agnostic defense against adversarial illusions. The code is available at https://github.com/fatemehakb/adversarial-illusions-mitigation.
title Breaking the Illusion: Consensus-Based Generative Mitigation of Adversarial Illusions in Multi-Modal Embeddings
topic Machine Learning
url https://arxiv.org/abs/2511.21893