Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
Fuente:
arXiv
Salvato in:
| Autori principali: | Hossain, Md Zarif, Imteaj, Ahmed |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
di: Hossain, Md Zarif, et al.
Pubblicazione: (2024)
di: Hossain, Md Zarif, et al.
Pubblicazione: (2024)
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
di: Rashid, Md Rafi Ur, et al.
Pubblicazione: (2026)
di: Rashid, Md Rafi Ur, et al.
Pubblicazione: (2026)
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
di: Patel, Het, et al.
Pubblicazione: (2025)
di: Patel, Het, et al.
Pubblicazione: (2025)
Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
di: Khanal, Aja, et al.
Pubblicazione: (2025)
di: Khanal, Aja, et al.
Pubblicazione: (2025)
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
di: Zhou, Andy, et al.
Pubblicazione: (2024)
di: Zhou, Andy, et al.
Pubblicazione: (2024)
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
di: Zhao, Yunhan, et al.
Pubblicazione: (2024)
di: Zhao, Yunhan, et al.
Pubblicazione: (2024)
White-box Multimodal Jailbreaks Against Large Vision-Language Models
di: Wang, Ruofan, et al.
Pubblicazione: (2024)
di: Wang, Ruofan, et al.
Pubblicazione: (2024)
Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
di: Hossain, Md. Iqbal, et al.
Pubblicazione: (2025)
di: Hossain, Md. Iqbal, et al.
Pubblicazione: (2025)
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025)
di: Liang, Shuang, et al.
Pubblicazione: (2025)
Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach
di: Guan, Jiwei, et al.
Pubblicazione: (2024)
di: Guan, Jiwei, et al.
Pubblicazione: (2024)
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
di: Wu, Sihao, et al.
Pubblicazione: (2025)
di: Wu, Sihao, et al.
Pubblicazione: (2025)
Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
di: Zhao, Yunhan, et al.
Pubblicazione: (2025)
di: Zhao, Yunhan, et al.
Pubblicazione: (2025)
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
di: Cui, Kaiyuan, et al.
Pubblicazione: (2026)
di: Cui, Kaiyuan, et al.
Pubblicazione: (2026)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
di: Wang, Lu, et al.
Pubblicazione: (2025)
di: Wang, Lu, et al.
Pubblicazione: (2025)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025)
di: Liang, Shuang, et al.
Pubblicazione: (2025)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
di: Yan, Yuping, et al.
Pubblicazione: (2025)
di: Yan, Yuping, et al.
Pubblicazione: (2025)
Jailbreaks on Vision Language Model via Multimodal Reasoning
di: Noheria, Aarush, et al.
Pubblicazione: (2026)
di: Noheria, Aarush, et al.
Pubblicazione: (2026)
Adversarial Attack Against Images Classification based on Generative Adversarial Networks
di: Yang, Yahe
Pubblicazione: (2024)
di: Yang, Yahe
Pubblicazione: (2024)
Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
di: Zhang, Naifu, et al.
Pubblicazione: (2025)
di: Zhang, Naifu, et al.
Pubblicazione: (2025)
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
di: Long, Jiahuan, et al.
Pubblicazione: (2025)
di: Long, Jiahuan, et al.
Pubblicazione: (2025)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
di: Tao, Xijia, et al.
Pubblicazione: (2024)
di: Tao, Xijia, et al.
Pubblicazione: (2024)
Brain Tumor Classifiers Under Attack: Robustness of ResNet Variants Against Transferable FGSM and PGD Attacks
di: Deem, Ryan, et al.
Pubblicazione: (2026)
di: Deem, Ryan, et al.
Pubblicazione: (2026)
Efficient Model-Based Purification Against Adversarial Attacks for LiDAR Segmentation
di: Gkillas, Alexandros, et al.
Pubblicazione: (2025)
di: Gkillas, Alexandros, et al.
Pubblicazione: (2025)
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
di: Hou, Jiacheng, et al.
Pubblicazione: (2026)
di: Hou, Jiacheng, et al.
Pubblicazione: (2026)
Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation
di: Ji, Yuheng, et al.
Pubblicazione: (2024)
di: Ji, Yuheng, et al.
Pubblicazione: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
di: Li, Jiayu, et al.
Pubblicazione: (2025)
di: Li, Jiayu, et al.
Pubblicazione: (2025)
A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models
di: Chen, Wutao, et al.
Pubblicazione: (2026)
di: Chen, Wutao, et al.
Pubblicazione: (2026)
HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models
di: Liu, Han, et al.
Pubblicazione: (2026)
di: Liu, Han, et al.
Pubblicazione: (2026)
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models
di: Panos, Aristeidis, et al.
Pubblicazione: (2024)
di: Panos, Aristeidis, et al.
Pubblicazione: (2024)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
di: Zheng, Weijie, et al.
Pubblicazione: (2024)
di: Zheng, Weijie, et al.
Pubblicazione: (2024)
IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves
di: Wang, Ruofan, et al.
Pubblicazione: (2024)
di: Wang, Ruofan, et al.
Pubblicazione: (2024)
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
di: Yin, Xiangyu, et al.
Pubblicazione: (2025)
di: Yin, Xiangyu, et al.
Pubblicazione: (2025)
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
di: Zhang, Jiaming, et al.
Pubblicazione: (2025)
di: Zhang, Jiaming, et al.
Pubblicazione: (2025)
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
di: Liu, Chaohu, et al.
Pubblicazione: (2025)
di: Liu, Chaohu, et al.
Pubblicazione: (2025)
TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
di: Li, Zhiwei, et al.
Pubblicazione: (2025)
di: Li, Zhiwei, et al.
Pubblicazione: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
di: Diao, Haiwen, et al.
Pubblicazione: (2025)
di: Diao, Haiwen, et al.
Pubblicazione: (2025)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
di: Wang, Yizhou, et al.
Pubblicazione: (2025)
Adversarial Prompt Distillation for Vision-Language Models
di: Luo, Lin, et al.
Pubblicazione: (2024)
di: Luo, Lin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
di: Hossain, Md Zarif, et al.
Pubblicazione: (2024) -
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
di: Rashid, Md Rafi Ur, et al.
Pubblicazione: (2026) -
Robust Vision-Language Models via Tensor Decomposition: A Defense Against Adversarial Attacks
di: Patel, Het, et al.
Pubblicazione: (2025) -
Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
di: Khanal, Aja, et al.
Pubblicazione: (2025) -
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
di: Zhou, Andy, et al.
Pubblicazione: (2024)