Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shen, Yao, Liuyi, Niu, Wujia, Zhang, Lan, Li, Yaliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safety Layers in Aligned Large Language Models: The Key to LLM Security
by: Li, Shen, et al.
Published: (2024)
by: Li, Shen, et al.
Published: (2024)
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024)
by: Feng, Qianhan, et al.
Published: (2024)
Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding
by: Li, Zhuoming, et al.
Published: (2025)
by: Li, Zhuoming, et al.
Published: (2025)
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
by: Xu, Shicheng, et al.
Published: (2024)
by: Xu, Shicheng, et al.
Published: (2024)
LVLM-Aided Alignment of Task-Specific Vision Models
by: Koebler, Alexander, et al.
Published: (2025)
by: Koebler, Alexander, et al.
Published: (2025)
GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models
by: Shen, Guanxi
Published: (2025)
by: Shen, Guanxi
Published: (2025)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
by: Jiao, Qirui, et al.
Published: (2025)
by: Jiao, Qirui, et al.
Published: (2025)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
by: Shao, Wenqi, et al.
Published: (2023)
by: Shao, Wenqi, et al.
Published: (2023)
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
by: Park, Seongheon, et al.
Published: (2026)
by: Park, Seongheon, et al.
Published: (2026)
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
by: Qharabagh, Muhammad Fetrat, et al.
Published: (2024)
by: Qharabagh, Muhammad Fetrat, et al.
Published: (2024)
Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation
by: Mao, Jiawei, et al.
Published: (2026)
by: Mao, Jiawei, et al.
Published: (2026)
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
by: Li, Yan, et al.
Published: (2024)
by: Li, Yan, et al.
Published: (2024)
Vision-and-Language Navigation via Causal Learning
by: Wang, Liuyi, et al.
Published: (2024)
by: Wang, Liuyi, et al.
Published: (2024)
SPARC: Concept-Aligned Sparse Autoencoders for Cross-Model and Cross-Modal Interpretability
by: Nasiri-Sarvi, Ali, et al.
Published: (2025)
by: Nasiri-Sarvi, Ali, et al.
Published: (2025)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
by: Shen, Zhixuan, et al.
Published: (2024)
by: Shen, Zhixuan, et al.
Published: (2024)
MAGIC: Meta-Ability Guided Interactive Chain-of-Distillation for Effective-and-Efficient Vision-and-Language Navigation
by: Wang, Liuyi, et al.
Published: (2024)
by: Wang, Liuyi, et al.
Published: (2024)
Adversarial Orthogonal Disentanglement for LVLM Hallucination Mitigation
by: Cheng, Ruoxi, et al.
Published: (2026)
by: Cheng, Ruoxi, et al.
Published: (2026)
Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
by: Yamabe, Shojiro, et al.
Published: (2025)
by: Yamabe, Shojiro, et al.
Published: (2025)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
by: Jiao, Qirui, et al.
Published: (2024)
by: Jiao, Qirui, et al.
Published: (2024)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
by: Kim, Mingyeong, et al.
Published: (2026)
by: Kim, Mingyeong, et al.
Published: (2026)
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction
by: Man, Yuanbin, et al.
Published: (2024)
by: Man, Yuanbin, et al.
Published: (2024)
Multi-Modal Character Localization and Extraction for Chinese Text Recognition
by: Li, Qilong, et al.
Published: (2026)
by: Li, Qilong, et al.
Published: (2026)
OmniCT: Towards a Unified Slice-Volume LVLM for Comprehensive CT Analysis
by: Lin, Tianwei, et al.
Published: (2026)
by: Lin, Tianwei, et al.
Published: (2026)
Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation
by: Wang, Jiaxi, et al.
Published: (2023)
by: Wang, Jiaxi, et al.
Published: (2023)
Seeing is Believing: Robust Vision-Guided Cross-Modal Prompt Learning under Label Noise
by: Geng, Zibin, et al.
Published: (2026)
by: Geng, Zibin, et al.
Published: (2026)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
by: Guo, Shasha, et al.
Published: (2025)
by: Guo, Shasha, et al.
Published: (2025)
AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrieval
by: Wang, Yihan, et al.
Published: (2026)
by: Wang, Yihan, et al.
Published: (2026)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Beyond Cross-Modal Alignment: Measuring and Leveraging Modality Gap in Vision-Language Models
by: Yan, Hanqi, et al.
Published: (2025)
by: Yan, Hanqi, et al.
Published: (2025)
REALM: An RGB and Event Aligned Latent Manifold for Cross-Modal Perception
by: Polizzi, Vincenzo, et al.
Published: (2026)
by: Polizzi, Vincenzo, et al.
Published: (2026)
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
by: Huang, Yiyang, et al.
Published: (2025)
by: Huang, Yiyang, et al.
Published: (2025)
Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning
by: Li, Jiazheng, et al.
Published: (2026)
by: Li, Jiazheng, et al.
Published: (2026)
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
by: Jiao, Qirui, et al.
Published: (2024)
by: Jiao, Qirui, et al.
Published: (2024)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
by: Miyamoto, Ryoto, et al.
Published: (2025)
by: Miyamoto, Ryoto, et al.
Published: (2025)
DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
by: Song, Jingyu, et al.
Published: (2025)
by: Song, Jingyu, et al.
Published: (2025)
Cross-Modal Adapter for Vision-Language Retrieval
by: Jiang, Haojun, et al.
Published: (2022)
by: Jiang, Haojun, et al.
Published: (2022)
Similar Items
-
Safety Layers in Aligned Large Language Models: The Key to LLM Security
by: Li, Shen, et al.
Published: (2024) -
Cross-Modal Attention Calibration for LVLM Hallucination Mitigation
by: Li, Jiaming, et al.
Published: (2025) -
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
by: Qin, Zhen, et al.
Published: (2024) -
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model
by: Feng, Qianhan, et al.
Published: (2024) -
Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding
by: Li, Zhuoming, et al.
Published: (2025)