Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
Fuente:
arXiv
Saved in:
| Main Authors: | Gou, Yunhao, Chen, Kai, Liu, Zhili, Hong, Lanqing, Xu, Hang, Li, Zhenguo, Yeung, Dit-Yan, Kwok, James T., Zhang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning
by: Gou, Yunhao, et al.
Published: (2023)
by: Gou, Yunhao, et al.
Published: (2023)
Mixed Autoencoder for Self-supervised Visual Representation Learning
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
Implicit Concept Removal of Diffusion Models
by: Liu, Zhili, et al.
Published: (2023)
by: Liu, Zhili, et al.
Published: (2023)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
MagicDrive: Street View Generation with Diverse 3D Geometry Control
by: Gao, Ruiyuan, et al.
Published: (2023)
by: Gao, Ruiyuan, et al.
Published: (2023)
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
by: Jiang, Chenhan, et al.
Published: (2025)
by: Jiang, Chenhan, et al.
Published: (2025)
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
by: Li, Pengxiang, et al.
Published: (2023)
by: Li, Pengxiang, et al.
Published: (2023)
TransformMix: Learning Transformation and Mixing Strategies from Data
by: Cheung, Tsz-Him, et al.
Published: (2024)
by: Cheung, Tsz-Him, et al.
Published: (2024)
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation
by: Jiang, Chenhan, et al.
Published: (2024)
by: Jiang, Chenhan, et al.
Published: (2024)
Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs
by: Zhong, Yingji, et al.
Published: (2024)
by: Zhong, Yingji, et al.
Published: (2024)
Learning High-resolution Vector Representation from Multi-Camera Images for 3D Object Detection
by: Chen, Zhili, et al.
Published: (2024)
by: Chen, Zhili, et al.
Published: (2024)
ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
IRWE: Inductive Random Walk for Joint Inference of Identity and Position Network Embedding
by: Qin, Meng, et al.
Published: (2024)
by: Qin, Meng, et al.
Published: (2024)
Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
by: Gao, Jiahui, et al.
Published: (2024)
by: Gao, Jiahui, et al.
Published: (2024)
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
by: Gao, Ruiyuan, et al.
Published: (2024)
by: Gao, Ruiyuan, et al.
Published: (2024)
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
by: Gao, Ruiyuan, et al.
Published: (2024)
by: Gao, Ruiyuan, et al.
Published: (2024)
Animate124: Animating One Image to 4D Dynamic Scene
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
by: Chen, Junsong, et al.
Published: (2023)
by: Chen, Junsong, et al.
Published: (2023)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
by: Tong, Tony Cheng, et al.
Published: (2024)
by: Tong, Tony Cheng, et al.
Published: (2024)
Multimodal Medical Image Binding via Shared Text Embeddings
by: Liu, Yunhao, et al.
Published: (2025)
by: Liu, Yunhao, et al.
Published: (2025)
DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
by: Chung, Tsz Ting, et al.
Published: (2025)
by: Chung, Tsz Ting, et al.
Published: (2025)
FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
by: Jiang, Chenhan, et al.
Published: (2026)
by: Jiang, Chenhan, et al.
Published: (2026)
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
Pre-train and Refine: Towards Higher Efficiency in K-Agnostic Community Detection without Quality Degradation
by: Qin, Meng, et al.
Published: (2024)
by: Qin, Meng, et al.
Published: (2024)
Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views
by: Zhong, Yingji, et al.
Published: (2025)
by: Zhong, Yingji, et al.
Published: (2025)
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
by: Zhang, Jiaming, et al.
Published: (2024)
by: Zhang, Jiaming, et al.
Published: (2024)
Safety of Multimodal Large Language Models on Images and Texts
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
Enhancing Text-to-Image Editing via Hybrid Mask-Informed Fusion
by: Li, Aoxue, et al.
Published: (2024)
by: Li, Aoxue, et al.
Published: (2024)
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
Similar Items
-
Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning
by: Gou, Yunhao, et al.
Published: (2023) -
Mixed Autoencoder for Self-supervised Visual Representation Learning
by: Chen, Kai, et al.
Published: (2023) -
Implicit Concept Removal of Diffusion Models
by: Liu, Zhili, et al.
Published: (2023) -
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025) -
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
by: Liu, Zhili, et al.
Published: (2024)