Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
Fuente:
arXiv
Saved in:
| Main Authors: | Shayegani, Erfan, Shahariar, G M, Abdali, Sara, Yu, Lei, Abu-Ghazaleh, Nael, Dong, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
That Doesn't Go There: Attacks on Shared State in Multi-User Augmented Reality Applications
by: Slocum, Carter, et al.
Published: (2023)
by: Slocum, Carter, et al.
Published: (2023)
Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems
by: Zhang, Yicheng, et al.
Published: (2025)
by: Zhang, Yicheng, et al.
Published: (2025)
Modeling Hierarchical Thinking in Large Reasoning Models
by: Shahariar, G M, et al.
Published: (2025)
by: Shahariar, G M, et al.
Published: (2025)
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
by: Abdali, Sara, et al.
Published: (2024)
by: Abdali, Sara, et al.
Published: (2024)
I Know What You Sync: Covert and Side Channel Attacks on File Systems via syncfs
by: Gu, Cheng, et al.
Published: (2024)
by: Gu, Cheng, et al.
Published: (2024)
Siren Song: Manipulating Pose Estimation in XR Headsets Using Acoustic Attacks
by: Huang, Zijian, et al.
Published: (2025)
by: Huang, Zijian, et al.
Published: (2025)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
Token-based Vehicular Security System (TVSS): Scalable, Secure, Low-latency Public Key Infrastructure for Connected Vehicles
by: Rabiah, Abdulrahman Bin, et al.
Published: (2024)
by: Rabiah, Abdulrahman Bin, et al.
Published: (2024)
NVBleed: Covert and Side-Channel Attacks on NVIDIA Multi-GPU Interconnect
by: Zhang, Yicheng, et al.
Published: (2025)
by: Zhang, Yicheng, et al.
Published: (2025)
Beyond the Bridge: Contention-Based Covert and Side Channel Attacks on Multi-GPU Interconnect
by: Zhang, Yicheng, et al.
Published: (2024)
by: Zhang, Yicheng, et al.
Published: (2024)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
by: Shahariar, G M, et al.
Published: (2024)
by: Shahariar, G M, et al.
Published: (2024)
ShadowScope: GPU Monitoring and Validation via Composable Side Channel Signals
by: Almusaddar, Ghadeer, et al.
Published: (2025)
by: Almusaddar, Ghadeer, et al.
Published: (2025)
Evaluating Granularity in Markov Chain-Based Trust Models for Vehicular Ad Hoc Networks (VANETs)
by: Shahariar, Rezvi
Published: (2026)
by: Shahariar, Rezvi
Published: (2026)
Securing the Future: Proactive Threat Hunting for Sustainable IoT Ecosystems
by: Ghasemshirazi, Saeid, et al.
Published: (2024)
by: Ghasemshirazi, Saeid, et al.
Published: (2024)
Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling
by: Li, Zida, et al.
Published: (2026)
by: Li, Zida, et al.
Published: (2026)
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
by: Xiong, Yuan, et al.
Published: (2025)
by: Xiong, Yuan, et al.
Published: (2025)
Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
by: Pan, Zihao, et al.
Published: (2025)
by: Pan, Zihao, et al.
Published: (2025)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents
by: Ding, Xuwei, et al.
Published: (2026)
by: Ding, Xuwei, et al.
Published: (2026)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
by: Bachu, Saketh, et al.
Published: (2024)
by: Bachu, Saketh, et al.
Published: (2024)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
by: Zhang, Rui, et al.
Published: (2026)
by: Zhang, Rui, et al.
Published: (2026)
The Need Of Trustworthy Announcements To Achieve Driving Comfort
by: Shahariar, Rezvi, et al.
Published: (2024)
by: Shahariar, Rezvi, et al.
Published: (2024)
A trust management framework for vehicular ad hoc networks
by: Shahariar, Rezvi, et al.
Published: (2024)
by: Shahariar, Rezvi, et al.
Published: (2024)
A fuzzy reward and punishment scheme for vehicular ad hoc networks
by: Shahariar, Rezvi, et al.
Published: (2024)
by: Shahariar, Rezvi, et al.
Published: (2024)
A Survey of Security Threats and Trust Management in Vehicular Ad Hoc Networks
by: Shahariar, Rezvi, et al.
Published: (2026)
by: Shahariar, Rezvi, et al.
Published: (2026)
Adversarial Examples are Misaligned in Diffusion Model Manifolds
by: Lorenz, Peter, et al.
Published: (2024)
by: Lorenz, Peter, et al.
Published: (2024)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
by: Sorokoletova, Olga E., et al.
Published: (2026)
by: Sorokoletova, Olga E., et al.
Published: (2026)
Exposing Citation Vulnerabilities in Generative Engines
by: Mochizuki, Riku, et al.
Published: (2025)
by: Mochizuki, Riku, et al.
Published: (2025)
Adversarial Text Generation with Dynamic Contextual Perturbation
by: Waghela, Hetvi, et al.
Published: (2025)
by: Waghela, Hetvi, et al.
Published: (2025)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
by: Liu, Lijia, et al.
Published: (2025)
by: Liu, Lijia, et al.
Published: (2025)
Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
by: Rad, Melissa Kazemi, et al.
Published: (2025)
by: Rad, Melissa Kazemi, et al.
Published: (2025)
Integrating Public Input and Technical Expertise for Effective Cybersecurity Policy Formulation
by: Ngobeni, Hlekane, et al.
Published: (2025)
by: Ngobeni, Hlekane, et al.
Published: (2025)
Watermarking Conditional Text Generation for AI Detection: Unveiling Challenges and a Semantic-Aware Watermark Remedy
by: Fu, Yu, et al.
Published: (2023)
by: Fu, Yu, et al.
Published: (2023)
SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
by: Lu, Weikai, et al.
Published: (2025)
by: Lu, Weikai, et al.
Published: (2025)
Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects
by: Rahman, Anichur, et al.
Published: (2025)
by: Rahman, Anichur, et al.
Published: (2025)
Similar Items
-
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025) -
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024) -
That Doesn't Go There: Attacks on Shared State in Multi-User Augmented Reality Applications
by: Slocum, Carter, et al.
Published: (2023) -
Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems
by: Zhang, Yicheng, et al.
Published: (2025) -
Modeling Hierarchical Thinking in Large Reasoning Models
by: Shahariar, G M, et al.
Published: (2025)