Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Ziwei, Zhao, Junyao, Yang, Le, He, Lijun, Li, Fan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Test-Time Attention Purification for Backdoored Large Vision Language Models
di: Zhang, Zhifang, et al.
Pubblicazione: (2026)
di: Zhang, Zhifang, et al.
Pubblicazione: (2026)
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
di: Ding, Yi, et al.
Pubblicazione: (2025)
di: Ding, Yi, et al.
Pubblicazione: (2025)
DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
di: Zhang, Yitong, et al.
Pubblicazione: (2025)
di: Zhang, Yitong, et al.
Pubblicazione: (2025)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
di: Ye, Mang, et al.
Pubblicazione: (2025)
di: Ye, Mang, et al.
Pubblicazione: (2025)
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
di: Sheng, Lijun, et al.
Pubblicazione: (2025)
di: Sheng, Lijun, et al.
Pubblicazione: (2025)
AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models
di: Li, Yue, et al.
Pubblicazione: (2026)
di: Li, Yue, et al.
Pubblicazione: (2026)
Revisiting Data Auditing in Large Vision-Language Models
di: Zhu, Hongyu, et al.
Pubblicazione: (2025)
di: Zhu, Hongyu, et al.
Pubblicazione: (2025)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
di: Gao, Kuofeng, et al.
Pubblicazione: (2024)
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
di: Yang, Hao, et al.
Pubblicazione: (2026)
di: Yang, Hao, et al.
Pubblicazione: (2026)
Image-Based Geolocation Using Large Vision-Language Models
di: Liu, Yi, et al.
Pubblicazione: (2024)
di: Liu, Yi, et al.
Pubblicazione: (2024)
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
di: Zhong, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zhong, Zhiyuan, et al.
Pubblicazione: (2025)
Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
di: Zhang, Rui, et al.
Pubblicazione: (2025)
di: Zhang, Rui, et al.
Pubblicazione: (2025)
Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security
di: Fan, Yihe, et al.
Pubblicazione: (2024)
di: Fan, Yihe, et al.
Pubblicazione: (2024)
Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models
di: Wu, Zongyu, et al.
Pubblicazione: (2025)
di: Wu, Zongyu, et al.
Pubblicazione: (2025)
Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory
di: Zhang, Ce, et al.
Pubblicazione: (2026)
di: Zhang, Ce, et al.
Pubblicazione: (2026)
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization
di: Hao, Shuyang, et al.
Pubblicazione: (2025)
di: Hao, Shuyang, et al.
Pubblicazione: (2025)
MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
di: Sun, Xiongtao, et al.
Pubblicazione: (2025)
di: Sun, Xiongtao, et al.
Pubblicazione: (2025)
Text is All You Need for Vision-Language Model Jailbreaking
di: Chen, Yihang, et al.
Pubblicazione: (2026)
di: Chen, Yihang, et al.
Pubblicazione: (2026)
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
di: Liu, Qin, et al.
Pubblicazione: (2025)
di: Liu, Qin, et al.
Pubblicazione: (2025)
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models
di: Yang, Ying, et al.
Pubblicazione: (2025)
di: Yang, Ying, et al.
Pubblicazione: (2025)
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
di: Wang, Yining, et al.
Pubblicazione: (2025)
di: Wang, Yining, et al.
Pubblicazione: (2025)
Similarity Distribution based Membership Inference Attack on Person Re-identification
di: Gao, Junyao, et al.
Pubblicazione: (2022)
di: Gao, Junyao, et al.
Pubblicazione: (2022)
Membership Inference Attack Against Masked Image Modeling
di: Li, Zheng, et al.
Pubblicazione: (2024)
di: Li, Zheng, et al.
Pubblicazione: (2024)
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
di: Wu, Jinman, et al.
Pubblicazione: (2026)
di: Wu, Jinman, et al.
Pubblicazione: (2026)
Jailbreak Vision Language Models via Bi-Modal Adversarial Prompt
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
di: Ying, Zonghao, et al.
Pubblicazione: (2024)
FIDAVL: Fake Image Detection and Attribution using Vision-Language Model
di: Keita, Mamadou, et al.
Pubblicazione: (2024)
di: Keita, Mamadou, et al.
Pubblicazione: (2024)
ViTGuard: Attention-aware Detection against Adversarial Examples for Vision Transformer
di: Sun, Shihua, et al.
Pubblicazione: (2024)
di: Sun, Shihua, et al.
Pubblicazione: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
di: Li, Jiayu, et al.
Pubblicazione: (2025)
di: Li, Jiayu, et al.
Pubblicazione: (2025)
Harnessing the Power of Large Vision Language Models for Synthetic Image Detection
di: Keita, Mamadou, et al.
Pubblicazione: (2024)
di: Keita, Mamadou, et al.
Pubblicazione: (2024)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
di: Liang, Shuang, et al.
Pubblicazione: (2025)
di: Liang, Shuang, et al.
Pubblicazione: (2025)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
di: Pan, Zihao, et al.
Pubblicazione: (2025)
di: Pan, Zihao, et al.
Pubblicazione: (2025)
ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
di: Cao, Hanwen, et al.
Pubblicazione: (2025)
di: Cao, Hanwen, et al.
Pubblicazione: (2025)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
di: Zhang, Xinwei, et al.
Pubblicazione: (2026)
di: Zhang, Xinwei, et al.
Pubblicazione: (2026)
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
di: Yin, Ziyi, et al.
Pubblicazione: (2023)
di: Yin, Ziyi, et al.
Pubblicazione: (2023)
PromptSmooth: Certifying Robustness of Medical Vision-Language Models via Prompt Learning
di: Hussein, Noor, et al.
Pubblicazione: (2024)
di: Hussein, Noor, et al.
Pubblicazione: (2024)
Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models
di: Guesmi, Amira, et al.
Pubblicazione: (2026)
di: Guesmi, Amira, et al.
Pubblicazione: (2026)
Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios
di: Qu, Jingen, et al.
Pubblicazione: (2025)
di: Qu, Jingen, et al.
Pubblicazione: (2025)
When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation Paradigm
di: Leng, Ye, et al.
Pubblicazione: (2026)
di: Leng, Ye, et al.
Pubblicazione: (2026)
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models
di: Fang, Hao, et al.
Pubblicazione: (2024)
di: Fang, Hao, et al.
Pubblicazione: (2024)
$\mathbf{S^2LM}$: Towards Semantic Steganography via Large Language Models
di: Wu, Huanqi, et al.
Pubblicazione: (2025)
di: Wu, Huanqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Test-Time Attention Purification for Backdoored Large Vision Language Models
di: Zhang, Zhifang, et al.
Pubblicazione: (2026) -
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
di: Ding, Yi, et al.
Pubblicazione: (2025) -
DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
di: Zhang, Yitong, et al.
Pubblicazione: (2025) -
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
di: Ye, Mang, et al.
Pubblicazione: (2025) -
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
di: Sheng, Lijun, et al.
Pubblicazione: (2025)