DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916780175261696 |
|---|---|
| author | Liu, Jianyu Guo, Hangyu Duan, Ranjie Bu, Xingyuan He, Yancheng Li, Shilong Huang, Hui Liu, Jiaheng Wang, Yucheng Jing, Chenchen Qu, Xingwei Zhang, Xiao Tan, Yingshui Wu, Yanan Gu, Jihao Li, Yangguang Zhu, Jianke |
| author_facet | Liu, Jianyu Guo, Hangyu Duan, Ranjie Bu, Xingyuan He, Yancheng Li, Shilong Huang, Hui Liu, Jiaheng Wang, Yucheng Jing, Chenchen Qu, Xingwei Zhang, Xiao Tan, Yingshui Wu, Yanan Gu, Jihao Li, Yangguang Zhu, Jianke |
| contents | Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data, thereby introducing new dimensions of potential attacks and complex risk combinations. In this paper, we begin with a detailed analysis aimed at disentangling risks through step-by-step reasoning within multimodal inputs. We find that systematic multimodal risk disentanglement substantially enhances the risk awareness of MLLMs. Via leveraging the strong discriminative abilities of multimodal risk disentanglement, we further introduce \textbf{DREAM} (\textit{\textbf{D}isentangling \textbf{R}isks to \textbf{E}nhance Safety \textbf{A}lignment in \textbf{M}LLMs}), a novel approach that enhances safety alignment in MLLMs through supervised fine-tuning and iterative Reinforcement Learning from AI Feedback (RLAIF). Experimental results show that DREAM significantly boosts safety during both inference and training phases without compromising performance on normal tasks (namely oversafety), achieving a 16.17\% improvement in the SIUO safe\&effective score compared to GPT-4V. The data and code are available at https://github.com/Kizna1ver/DREAM. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_18053 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models Liu, Jianyu Guo, Hangyu Duan, Ranjie Bu, Xingyuan He, Yancheng Li, Shilong Huang, Hui Liu, Jiaheng Wang, Yucheng Jing, Chenchen Qu, Xingwei Zhang, Xiao Tan, Yingshui Wu, Yanan Gu, Jihao Li, Yangguang Zhu, Jianke Computation and Language Computer Vision and Pattern Recognition Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data, thereby introducing new dimensions of potential attacks and complex risk combinations. In this paper, we begin with a detailed analysis aimed at disentangling risks through step-by-step reasoning within multimodal inputs. We find that systematic multimodal risk disentanglement substantially enhances the risk awareness of MLLMs. Via leveraging the strong discriminative abilities of multimodal risk disentanglement, we further introduce \textbf{DREAM} (\textit{\textbf{D}isentangling \textbf{R}isks to \textbf{E}nhance Safety \textbf{A}lignment in \textbf{M}LLMs}), a novel approach that enhances safety alignment in MLLMs through supervised fine-tuning and iterative Reinforcement Learning from AI Feedback (RLAIF). Experimental results show that DREAM significantly boosts safety during both inference and training phases without compromising performance on normal tasks (namely oversafety), achieving a 16.17\% improvement in the SIUO safe\&effective score compared to GPT-4V. The data and code are available at https://github.com/Kizna1ver/DREAM. |
| title | DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models |
| topic | Computation and Language Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2504.18053 |