Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hao, Shuyang, Hooi, Bryan, Liu, Jun, Chang, Kai-Wei, Huang, Zi, Cai, Yujun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
How does Watermarking Affect Visual Language Models in Document Understanding?
von: Xu, Chunxue, et al.
Veröffentlicht: (2025)
von: Xu, Chunxue, et al.
Veröffentlicht: (2025)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation
von: Mia, Md Jueal, et al.
Veröffentlicht: (2025)
von: Mia, Md Jueal, et al.
Veröffentlicht: (2025)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2025)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2025)
Words or Vision: Do Vision-Language Models Have Blind Faith in Text?
von: Deng, Ailin, et al.
Veröffentlicht: (2025)
von: Deng, Ailin, et al.
Veröffentlicht: (2025)
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
von: Li, Sifan, et al.
Veröffentlicht: (2025)
von: Li, Sifan, et al.
Veröffentlicht: (2025)
Lost in Edits? A $λ$-Compass for AIGC Provenance
von: You, Wenhao, et al.
Veröffentlicht: (2025)
von: You, Wenhao, et al.
Veröffentlicht: (2025)
Jailbreaking Vision-Language Models Through the Visual Modality
von: Azulay, Aharon, et al.
Veröffentlicht: (2026)
von: Azulay, Aharon, et al.
Veröffentlicht: (2026)
Revealing Physical-World Semantic Vulnerabilities: Universal Adversarial Patches for Infrared Vision-Language Models
von: Hu, Chengyin, et al.
Veröffentlicht: (2026)
von: Hu, Chengyin, et al.
Veröffentlicht: (2026)
GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning
von: Li, Yujun, et al.
Veröffentlicht: (2026)
von: Li, Yujun, et al.
Veröffentlicht: (2026)
Jailbreak Large Vision-Language Models Through Multi-Modal Linkage
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
Jailbreaks on Vision Language Model via Multimodal Reasoning
von: Noheria, Aarush, et al.
Veröffentlicht: (2026)
von: Noheria, Aarush, et al.
Veröffentlicht: (2026)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
von: Wu, Sihao, et al.
Veröffentlicht: (2025)
von: Wu, Sihao, et al.
Veröffentlicht: (2025)
When Text and Images Don't Mix: Bias-Correcting Language-Image Similarity Scores for Anomaly Detection
von: Goodge, Adam, et al.
Veröffentlicht: (2024)
von: Goodge, Adam, et al.
Veröffentlicht: (2024)
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
Adversarial Prompt Distillation for Vision-Language Models
von: Luo, Lin, et al.
Veröffentlicht: (2024)
von: Luo, Lin, et al.
Veröffentlicht: (2024)
JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
von: Jin, Haibo, et al.
Veröffentlicht: (2024)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
von: Li, Sifan, et al.
Veröffentlicht: (2025)
von: Li, Sifan, et al.
Veröffentlicht: (2025)
Dataset Distillation via Vision-Language Category Prototype
von: Zou, Yawen, et al.
Veröffentlicht: (2025)
von: Zou, Yawen, et al.
Veröffentlicht: (2025)
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
von: Zhao, Shiji, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
von: Lu, Fan, et al.
Veröffentlicht: (2024)
von: Lu, Fan, et al.
Veröffentlicht: (2024)
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
von: Wu, Fangyu, et al.
Veröffentlicht: (2025)
von: Wu, Fangyu, et al.
Veröffentlicht: (2025)
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
von: Goodge, Adam, et al.
Veröffentlicht: (2025)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
LLMs are Good Action Recognizers
von: Qu, Haoxuan, et al.
Veröffentlicht: (2024)
von: Qu, Haoxuan, et al.
Veröffentlicht: (2024)
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
von: Liu, Qinying, et al.
Veröffentlicht: (2023)
von: Liu, Qinying, et al.
Veröffentlicht: (2023)
MMT-ARD: Multimodal Multi-Teacher Adversarial Distillation for Robust Vision-Language Models
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
When Lighting Deceives: Exposing Vision-Language Models' Illumination Vulnerability Through Illumination Transformation Attack
von: Liu, Hanqing, et al.
Veröffentlicht: (2025)
von: Liu, Hanqing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization
von: Hao, Shuyang, et al.
Veröffentlicht: (2025) -
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
von: Hao, Shuyang, et al.
Veröffentlicht: (2025) -
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025) -
How does Watermarking Affect Visual Language Models in Document Understanding?
von: Xu, Chunxue, et al.
Veröffentlicht: (2025) -
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
von: Ying, Zonghao, et al.
Veröffentlicht: (2024)