Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Bachu, Saketh, Shayegani, Erfan, Lal, Rohit, Chakraborty, Trishna, Dutta, Arindam, Song, Chengyu, Dong, Yue, Abu-Ghazaleh, Nael, Roy-Chowdhury, Amit K. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024)
by: Chakraborty, Trishna, et al.
Published: (2024)
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Modeling Hierarchical Thinking in Large Reasoning Models
by: Shahariar, G M, et al.
Published: (2025)
by: Shahariar, G M, et al.
Published: (2025)
VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation under real Occlusions
by: Garg, Yash, et al.
Published: (2025)
by: Garg, Yash, et al.
Published: (2025)
Unsupervised Domain Adaptation for Occlusion Resilient Human Pose Estimation
by: Dutta, Arindam, et al.
Published: (2025)
by: Dutta, Arindam, et al.
Published: (2025)
STRIDE: Single-video based Temporally Continuous Occlusion-Robust 3D Pose Estimation
by: Lal, Rohit, et al.
Published: (2023)
by: Lal, Rohit, et al.
Published: (2023)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
by: Shahgir, Haz Sameen, et al.
Published: (2026)
by: Shahgir, Haz Sameen, et al.
Published: (2026)
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
by: Mamun, Md Abdullah Al, et al.
Published: (2025)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Repurposing SAM for User-Defined Semantics Aware Segmentation
by: Kundu, Rohit, et al.
Published: (2023)
by: Kundu, Rohit, et al.
Published: (2023)
Multi-modal Pose Diffuser: A Multimodal Generative Conditional Pose Prior
by: Ta, Calvin-Khang, et al.
Published: (2024)
by: Ta, Calvin-Khang, et al.
Published: (2024)
That Doesn't Go There: Attacks on Shared State in Multi-User Augmented Reality Applications
by: Slocum, Carter, et al.
Published: (2023)
by: Slocum, Carter, et al.
Published: (2023)
Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems
by: Zhang, Yicheng, et al.
Published: (2025)
by: Zhang, Yicheng, et al.
Published: (2025)
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
by: Joshi, Maithili, et al.
Published: (2025)
by: Joshi, Maithili, et al.
Published: (2025)
Toward Autonomous Laboratory Safety Monitoring with Vision Language Models: Learning to See Hazards Through Scene Structure
by: Chakraborty, Trishna, et al.
Published: (2026)
by: Chakraborty, Trishna, et al.
Published: (2026)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
POSTURE: Pose Guided Unsupervised Domain Adaptation for Human Body Part Segmentation
by: Dutta, Arindam, et al.
Published: (2024)
by: Dutta, Arindam, et al.
Published: (2024)
Layer-wise Swapping for Generalizable Multilingual Safety
by: Shin, Hyunseo, et al.
Published: (2026)
by: Shin, Hyunseo, et al.
Published: (2026)
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
by: Ding, Tianye, et al.
Published: (2025)
by: Ding, Tianye, et al.
Published: (2025)
Co(ve)rtex: ML Models as storage channels and their (mis-)applications
by: Mamun, Md Abdullah Al, et al.
Published: (2023)
by: Mamun, Md Abdullah Al, et al.
Published: (2023)
Hippocampal Atrophy Patterns Across the Alzheimer's Disease Spectrum: A Voxel-Based Morphometry Analysis
by: Niraula, Trishna
Published: (2026)
by: Niraula, Trishna
Published: (2026)
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
by: Zaree, Pedram, et al.
Published: (2025)
by: Zaree, Pedram, et al.
Published: (2025)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
by: Vashishtha, Aniket, et al.
Published: (2023)
by: Vashishtha, Aniket, et al.
Published: (2023)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
by: Das, Arion, et al.
Published: (2026)
by: Das, Arion, et al.
Published: (2026)
From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations
by: Borah, Abhilekh, et al.
Published: (2025)
by: Borah, Abhilekh, et al.
Published: (2025)
Understanding Layer Significance in LLM Alignment
by: Shi, Guangyuan, et al.
Published: (2024)
by: Shi, Guangyuan, et al.
Published: (2024)
Layer-wise Representation Dynamics: An Empirical Investigation Across Embedders and Base LLMs
by: Jiang, Jingzhou, et al.
Published: (2026)
by: Jiang, Jingzhou, et al.
Published: (2026)
Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
by: Xu, Jingjing, et al.
Published: (2024)
by: Xu, Jingjing, et al.
Published: (2024)
Decision-Aware Trust Signal Alignment for SOC Alert Triage
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
by: Chowdhury, Israt Jahan, et al.
Published: (2026)
Alignment Adapter to Improve the Performance of Compressed Deep Learning Models
by: Rai, Rohit Raj, et al.
Published: (2026)
by: Rai, Rohit Raj, et al.
Published: (2026)
Rethinking the adaptive relationship between Encoder Layers and Decoder Layers
by: Song, Yubo
Published: (2024)
by: Song, Yubo
Published: (2024)
Attend First, Consolidate Later: On the Importance of Attention in Different LLM Layers
by: Ben-Artzy, Amit, et al.
Published: (2024)
by: Ben-Artzy, Amit, et al.
Published: (2024)
Alignment is Localized: A Causal Probe into Preference Layers
by: Chaudhury, Archie
Published: (2025)
by: Chaudhury, Archie
Published: (2025)
Leveraging Synthetic Adult Datasets for Unsupervised Infant Pose Estimation
by: Bose, Sarosij, et al.
Published: (2025)
by: Bose, Sarosij, et al.
Published: (2025)
LayerPeeler: Autoregressive Peeling for Layer-wise Image Vectorization
by: Wu, Ronghuan, et al.
Published: (2025)
by: Wu, Ronghuan, et al.
Published: (2025)
TRepLiNa: Layer-wise CKA+REPINA Alignment Improves Low-Resource Machine Translation in Aya-23 8B
by: Nakai, Toshiki, et al.
Published: (2025)
by: Nakai, Toshiki, et al.
Published: (2025)
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
by: Bose, Sarosij, et al.
Published: (2025)
by: Bose, Sarosij, et al.
Published: (2025)
Dual Alignment Between Language Model Layers and Human Sentence Processing
by: Kuribayashi, Tatsuki, et al.
Published: (2026)
by: Kuribayashi, Tatsuki, et al.
Published: (2026)
Uncertainty-Aware Diffusion Guided Refinement of 3D Scenes
by: Bose, Sarosij, et al.
Published: (2025)
by: Bose, Sarosij, et al.
Published: (2025)
Similar Items
-
Cross-Modal Safety Alignment: Is textual unlearning all you need?
by: Chakraborty, Trishna, et al.
Published: (2024) -
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
by: Shayegani, Erfan, et al.
Published: (2025) -
Modeling Hierarchical Thinking in Large Reasoning Models
by: Shahariar, G M, et al.
Published: (2025) -
VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation under real Occlusions
by: Garg, Yash, et al.
Published: (2025) -
Unsupervised Domain Adaptation for Occlusion Resilient Human Pose Estimation
by: Dutta, Arindam, et al.
Published: (2025)