VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Qin, Wang, Fei, Xiao, Chaowei, Chen, Muhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
FaceSwapGuard: Safeguarding Facial Privacy from DeepFake Threats through Identity Obfuscation
von: Wang, Li, et al.
Veröffentlicht: (2025)
von: Wang, Li, et al.
Veröffentlicht: (2025)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
von: Xiang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Xiang, Yuxiao, et al.
Veröffentlicht: (2025)
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
von: Xu, Peiyang, et al.
Veröffentlicht: (2025)
von: Xu, Peiyang, et al.
Veröffentlicht: (2025)
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2025)
IdentityGuard: Context-Aware Restriction and Provenance for Personalized Synthesis
von: Zhang, Lingyun, et al.
Veröffentlicht: (2026)
von: Zhang, Lingyun, et al.
Veröffentlicht: (2026)
PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
von: Li, Nanxi, et al.
Veröffentlicht: (2025)
von: Li, Nanxi, et al.
Veröffentlicht: (2025)
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
von: Yuan, Lingzhi, et al.
Veröffentlicht: (2025)
von: Yuan, Lingzhi, et al.
Veröffentlicht: (2025)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
Text is All You Need for Vision-Language Model Jailbreaking
von: Chen, Yihang, et al.
Veröffentlicht: (2026)
von: Chen, Yihang, et al.
Veröffentlicht: (2026)
Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
von: Xie, Chunlong, et al.
Veröffentlicht: (2025)
von: Xie, Chunlong, et al.
Veröffentlicht: (2025)
HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios
von: Pu, Jiayue, et al.
Veröffentlicht: (2026)
von: Pu, Jiayue, et al.
Veröffentlicht: (2026)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
von: Tang, Peiyuan, et al.
Veröffentlicht: (2025)
von: Tang, Peiyuan, et al.
Veröffentlicht: (2025)
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
von: Fan, Yucheng, et al.
Veröffentlicht: (2025)
DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
von: Zhang, Yitong, et al.
Veröffentlicht: (2025)
von: Zhang, Yitong, et al.
Veröffentlicht: (2025)
Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models
von: Liang, Shuang, et al.
Veröffentlicht: (2025)
von: Liang, Shuang, et al.
Veröffentlicht: (2025)
AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models
von: Li, Yue, et al.
Veröffentlicht: (2026)
von: Li, Yue, et al.
Veröffentlicht: (2026)
From Shortcuts to Triggers: Backdoor Defense with Denoised PoE
von: Liu, Qin, et al.
Veröffentlicht: (2023)
von: Liu, Qin, et al.
Veröffentlicht: (2023)
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyu, et al.
Veröffentlicht: (2025)
Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification
von: Meng, Xiangtao, et al.
Veröffentlicht: (2025)
von: Meng, Xiangtao, et al.
Veröffentlicht: (2025)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
von: Li, Jiayu, et al.
Veröffentlicht: (2025)
FreezeAsGuard: Mitigating Illegal Adaptation of Diffusion Models via Selective Tensor Freezing
von: Huang, Kai, et al.
Veröffentlicht: (2024)
von: Huang, Kai, et al.
Veröffentlicht: (2024)
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language Models
von: Zheng, Ziwei, et al.
Veröffentlicht: (2025)
von: Zheng, Ziwei, et al.
Veröffentlicht: (2025)
Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment
von: Wang, Li, et al.
Veröffentlicht: (2025)
von: Wang, Li, et al.
Veröffentlicht: (2025)
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
von: Ma, Bo, et al.
Veröffentlicht: (2026)
von: Ma, Bo, et al.
Veröffentlicht: (2026)
Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
von: Amebley, David, et al.
Veröffentlicht: (2025)
von: Amebley, David, et al.
Veröffentlicht: (2025)
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
von: Ma, Xingjun, et al.
Veröffentlicht: (2025)
von: Ma, Xingjun, et al.
Veröffentlicht: (2025)
Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
von: Li, Xiao, et al.
Veröffentlicht: (2026)
von: Li, Xiao, et al.
Veröffentlicht: (2026)
PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility
von: Shahariar, G M, et al.
Veröffentlicht: (2026)
von: Shahariar, G M, et al.
Veröffentlicht: (2026)
Adversarial Robustness of Vision in Open Foundation Models
von: Fox, Jonathon, et al.
Veröffentlicht: (2025)
von: Fox, Jonathon, et al.
Veröffentlicht: (2025)
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV Equalization
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
von: Hao, Shuyang, et al.
Veröffentlicht: (2025)
SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models
von: Alam, Mohammed Talha, et al.
Veröffentlicht: (2025)
von: Alam, Mohammed Talha, et al.
Veröffentlicht: (2025)
EO-VLM: VLM-Guided Energy Overload Attacks on Vision Models
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
von: Seo, Minjae, et al.
Veröffentlicht: (2025)
The Structural Safety Generalization Problem
von: Broomfield, Julius, et al.
Veröffentlicht: (2025)
von: Broomfield, Julius, et al.
Veröffentlicht: (2025)
PLA: Prompt Learning Attack against Text-to-Image Generative Models
von: Lyu, Xinqi, et al.
Veröffentlicht: (2025)
von: Lyu, Xinqi, et al.
Veröffentlicht: (2025)
Backdoor Attack Against Vision Transformers via Attention Gradient-Based Image Erosion
von: Guo, Ji, et al.
Veröffentlicht: (2024)
von: Guo, Ji, et al.
Veröffentlicht: (2024)
DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling
von: Li, Boheng, et al.
Veröffentlicht: (2025)
von: Li, Boheng, et al.
Veröffentlicht: (2025)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Fight Perturbations with Perturbations: Defending Adversarial Attacks via Neuron Influence
von: Chen, Ruoxi, et al.
Veröffentlicht: (2021)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
von: Wang, Yu, et al.
Veröffentlicht: (2024) -
FaceSwapGuard: Safeguarding Facial Privacy from DeepFake Threats through Identity Obfuscation
von: Wang, Li, et al.
Veröffentlicht: (2025) -
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
von: Xiang, Yuxiao, et al.
Veröffentlicht: (2025) -
SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
von: Xu, Peiyang, et al.
Veröffentlicht: (2025) -
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
von: Jiang, Zhengyuan, et al.
Veröffentlicht: (2025)