Unraveling and Mitigating Safety Alignment Degradation of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Qin, Shang, Chao, Liu, Ling, Pappas, Nikolaos, Ma, Jie, John, Neha Anna, Doss, Srikanth, Marquez, Lluis, Ballesteros, Miguel, Benajiba, Yassine |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference time LLM alignment in single and multidomain preference spectrum
by: Shahriar, Sadat, et al.
Published: (2024)
by: Shahriar, Sadat, et al.
Published: (2024)
Diable: Efficient Dialogue State Tracking as Operations on Tables
by: Lesci, Pietro, et al.
Published: (2023)
by: Lesci, Pietro, et al.
Published: (2023)
Active Evaluation Acquisition for Efficient LLM Benchmarking
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Arabic Named Entity Recognition
by: Yassine Benajiba
Published: (2010)
by: Yassine Benajiba
Published: (2010)
Towards Long Context Hallucination Detection
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
General Purpose Verification for Chain of Thought Prompting
by: Vacareanu, Robert, et al.
Published: (2024)
by: Vacareanu, Robert, et al.
Published: (2024)
Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
by: Qi, Zheng, et al.
Published: (2025)
by: Qi, Zheng, et al.
Published: (2025)
Balancing Classification and Calibration Performance in Decision-Making LLMs via Calibration Aware Reinforcement Learning
by: Yaldiz, Duygu Nur, et al.
Published: (2026)
by: Yaldiz, Duygu Nur, et al.
Published: (2026)
From Instructions to Constraints: Language Model Alignment with Automatic Constraint Verification
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Rethinking LLM Uncertainty: A Multi-Agent Approach to Estimating Black-Box Model Uncertainty
by: Feng, Yu, et al.
Published: (2024)
by: Feng, Yu, et al.
Published: (2024)
NewsQs: Multi-Source Question Generation for the Inquiring Mind
by: Hwang, Alyssa, et al.
Published: (2024)
by: Hwang, Alyssa, et al.
Published: (2024)
Sequential Editing for Lifelong Training of Speech Recognition Models
by: Kulshreshtha, Devang, et al.
Published: (2024)
by: Kulshreshtha, Devang, et al.
Published: (2024)
MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
by: Singh, Jyotika, et al.
Published: (2026)
by: Singh, Jyotika, et al.
Published: (2026)
Barriers to Discrete Reasoning with Transformers: A Survey Across Depth, Exactness, and Bandwidth
by: Yuan, Michelle, et al.
Published: (2026)
by: Yuan, Michelle, et al.
Published: (2026)
SQLSpace: A Representation Space for Text-to-SQL to Discover and Mitigate Robustness Gaps
by: Srikanth, Neha, et al.
Published: (2025)
by: Srikanth, Neha, et al.
Published: (2025)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
by: Srikanth, Neha, et al.
Published: (2025)
by: Srikanth, Neha, et al.
Published: (2025)
VLM-Guard: Safeguarding Vision-Language Models via Fulfilling Safety Alignment Gap
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
MemInsight: Autonomous Memory Augmentation for LLM Agents
by: Salama, Rana, et al.
Published: (2025)
by: Salama, Rana, et al.
Published: (2025)
Open Domain Question Answering with Conflicting Contexts
by: Liu, Siyi, et al.
Published: (2024)
by: Liu, Siyi, et al.
Published: (2024)
Safety Alignment for Vision Language Models
by: Liu, Zhendong, et al.
Published: (2024)
by: Liu, Zhendong, et al.
Published: (2024)
Exploiting Data Significance in Remote Estimation of Discrete-State Markov Sources
by: Luo, Jiping, et al.
Published: (2024)
by: Luo, Jiping, et al.
Published: (2024)
From AoI to QVAoI: Query-Based Semantics-Aware Scheduling for Energy-Harvesting IoT Systems
by: Delfani, Erfan, et al.
Published: (2024)
by: Delfani, Erfan, et al.
Published: (2024)
From Timestamps to Versions: Version AoI in Single- and Multi-Hop Networks
by: Delfani, Erfan, et al.
Published: (2025)
by: Delfani, Erfan, et al.
Published: (2025)
Real-Time Reconstruction and Actuation Error Analysis for Markov Sources over MPR Channels
by: Elessawy, Pansee S., et al.
Published: (2026)
by: Elessawy, Pansee S., et al.
Published: (2026)
Goal-oriented Estimation of Multiple Markov Sources in Resource-constrained Systems
by: Luo, Jiping, et al.
Published: (2023)
by: Luo, Jiping, et al.
Published: (2023)
Semantics-Aware Updates from Remote Energy Harvesting Devices to Interconnected LEO Satellites
by: Delfani, Erfan, et al.
Published: (2025)
by: Delfani, Erfan, et al.
Published: (2025)
Semantic-Aware Remote Estimation of Multiple Markov Sources Under Constraints
by: Luo, Jiping, et al.
Published: (2024)
by: Luo, Jiping, et al.
Published: (2024)
Computation-aware Energy-harvesting Federated Learning: Cyclic Scheduling with Selective Participation
by: Jeong, Eunjeong, et al.
Published: (2025)
by: Jeong, Eunjeong, et al.
Published: (2025)
Battery-aware Cyclic Scheduling in Energy-harvesting Federated Learning
by: Jeong, Eunjeong, et al.
Published: (2025)
by: Jeong, Eunjeong, et al.
Published: (2025)
On the Role of Age and Semantics of Information in Remote Estimation of Markov Sources
by: Luo, Jiping, et al.
Published: (2025)
by: Luo, Jiping, et al.
Published: (2025)
Technology‐Enabled Competitiveness and Experiences in Tourism: A Transformative Era
by: Eleni Michopoulou, et al.
Published: (2025)
by: Eleni Michopoulou, et al.
Published: (2025)
Computing the Exact Pareto Front in Average-Cost Multi-Objective Markov Decision Processes
by: Luo, Jiping, et al.
Published: (2026)
by: Luo, Jiping, et al.
Published: (2026)
Optimizing Version AoI in Energy-Harvesting IoT: Model-Based and Learning-Based Approaches
by: Delfani, Erfan, et al.
Published: (2025)
by: Delfani, Erfan, et al.
Published: (2025)
On the Cost of Consecutive Estimation Error: Significance-Aware Non-linear Aging
by: Luo, Jiping, et al.
Published: (2024)
by: Luo, Jiping, et al.
Published: (2024)
Joint Accuracy and Confidentiality in Semantic-Aware Secure Remote Reconstruction
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
Optimizing Information Freshness in IoT Systems with Update Rate Constraints: A Token-Based Approach
by: Delfani, Erfan, et al.
Published: (2024)
by: Delfani, Erfan, et al.
Published: (2024)
Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment
by: Xue, Zhiyu, et al.
Published: (2026)
by: Xue, Zhiyu, et al.
Published: (2026)
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
by: Zhou, Zhanhui, et al.
Published: (2024)
by: Zhou, Zhanhui, et al.
Published: (2024)
Understanding and Improving Information Preservation in Prompt Compression for LLMs
by: Łajewska, Weronika, et al.
Published: (2025)
by: Łajewska, Weronika, et al.
Published: (2025)
Watermarking Degrades Alignment in Language Models: Analysis and Mitigation
by: Verma, Apurv, et al.
Published: (2025)
by: Verma, Apurv, et al.
Published: (2025)
Similar Items
-
Inference time LLM alignment in single and multidomain preference spectrum
by: Shahriar, Sadat, et al.
Published: (2024) -
Diable: Efficient Dialogue State Tracking as Operations on Tables
by: Lesci, Pietro, et al.
Published: (2023) -
Active Evaluation Acquisition for Efficient LLM Benchmarking
by: Li, Yang, et al.
Published: (2024) -
Arabic Named Entity Recognition
by: Yassine Benajiba
Published: (2010) -
Towards Long Context Hallucination Detection
by: Liu, Siyi, et al.
Published: (2025)