Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Enyi, Shen, Fei, Miao, Shuyi, Zhu, Linxia, Shao, Pengyang, Tang, Jinhui, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
von: Shi, Enyi, et al.
Veröffentlicht: (2026)
von: Shi, Enyi, et al.
Veröffentlicht: (2026)
Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons
von: Zhang, Xianhui, et al.
Veröffentlicht: (2026)
von: Zhang, Xianhui, et al.
Veröffentlicht: (2026)
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
von: Wang, Zhaoxin, et al.
Veröffentlicht: (2026)
von: Wang, Zhaoxin, et al.
Veröffentlicht: (2026)
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026)
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
von: Shen, Fei, et al.
Veröffentlicht: (2025)
von: Shen, Fei, et al.
Veröffentlicht: (2025)
MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval
von: Zhu, Fengbin, et al.
Veröffentlicht: (2026)
von: Zhu, Fengbin, et al.
Veröffentlicht: (2026)
UniDetect: LLM-Driven Universal Fraud Detection across Heterogeneous Blockchains
von: Miao, Shuyi, et al.
Veröffentlicht: (2026)
von: Miao, Shuyi, et al.
Veröffentlicht: (2026)
DRAFT: Task Decoupled Latent Reasoning for Agent Safety
von: Wang, Lin, et al.
Veröffentlicht: (2026)
von: Wang, Lin, et al.
Veröffentlicht: (2026)
Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models
von: Li, Shaotian, et al.
Veröffentlicht: (2026)
von: Li, Shaotian, et al.
Veröffentlicht: (2026)
Universal Scene Graph Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
Controllable Value Alignment in Large Language Models through Neuron-Level Editing
von: Yang, Yonghui, et al.
Veröffentlicht: (2026)
von: Yang, Yonghui, et al.
Veröffentlicht: (2026)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
von: Luo, Jiani, et al.
Veröffentlicht: (2026)
von: Luo, Jiani, et al.
Veröffentlicht: (2026)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities
von: Dou, Jingtong, et al.
Veröffentlicht: (2026)
von: Dou, Jingtong, et al.
Veröffentlicht: (2026)
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
XNLP: An Interactive Demonstration System for Universal Structured NLP
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
NExT-Guard: Training-Free Streaming Safeguard without Token-Level Labels
von: Fang, Junfeng, et al.
Veröffentlicht: (2026)
von: Fang, Junfeng, et al.
Veröffentlicht: (2026)
AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
von: Guo, Yangyang, et al.
Veröffentlicht: (2024)
Learning to Ask Critical Questions for Assisting Product Search
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective
von: Zheng, Zhedong, et al.
Veröffentlicht: (2022)
von: Zheng, Zhedong, et al.
Veröffentlicht: (2022)
ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
von: Khalil, Ahmad, et al.
Veröffentlicht: (2025)
von: Khalil, Ahmad, et al.
Veröffentlicht: (2025)
NExT-GPT: Any-to-Any Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2023)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2023)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
Do LLMs and VLMs Share Neurons for Inference? Evidence and Mechanisms of Cross-Modal Transfer
von: Cui, Chenhang, et al.
Veröffentlicht: (2026)
von: Cui, Chenhang, et al.
Veröffentlicht: (2026)
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
von: Huang, Youcheng, et al.
Veröffentlicht: (2024)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
von: Deng, Yang, et al.
Veröffentlicht: (2023)
von: Deng, Yang, et al.
Veröffentlicht: (2023)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
von: Zhang, An, et al.
Veröffentlicht: (2024)
von: Zhang, An, et al.
Veröffentlicht: (2024)
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
von: Fang, Junfeng, et al.
Veröffentlicht: (2025)
von: Fang, Junfeng, et al.
Veröffentlicht: (2025)
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
von: Ye, Ziang, et al.
Veröffentlicht: (2025)
von: Ye, Ziang, et al.
Veröffentlicht: (2025)
Search-in-the-Chain: Interactively Enhancing Large Language Models with Search for Knowledge-intensive Tasks
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
von: Xu, Shicheng, et al.
Veröffentlicht: (2023)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
IGD: Token Decisiveness Modeling via Information Gain in LLMs for Personalized Recommendation
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
ACE-Align: Attribute Causal Effect Alignment for Cultural Values under Varying Persona Granularities
von: Luo, Jiatang, et al.
Veröffentlicht: (2026)
von: Luo, Jiatang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models
von: Shi, Enyi, et al.
Veröffentlicht: (2026) -
Who Transfers Safety? Identifying and Targeting Cross-Lingual Shared Safety Neurons
von: Zhang, Xianhui, et al.
Veröffentlicht: (2026) -
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
von: Wang, Zhaoxin, et al.
Veröffentlicht: (2026) -
TraceRouter: Robust Safety for Large Foundation Models via Path-Level Intervention
von: Shi, Chuancheng, et al.
Veröffentlicht: (2026) -
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
von: Shen, Fei, et al.
Veröffentlicht: (2025)