ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cui, Shiyao, Zhang, Qinglin, Ouyang, Xuan, Chen, Renmiao, Zhang, Zhexin, Lu, Yida, Wang, Hongning, Qiu, Han, Huang, Minlie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910956874891264
author Cui, Shiyao
Zhang, Qinglin
Ouyang, Xuan
Chen, Renmiao
Zhang, Zhexin
Lu, Yida
Wang, Hongning
Qiu, Han
Huang, Minlie
author_facet Cui, Shiyao
Zhang, Qinglin
Ouyang, Xuan
Chen, Renmiao
Zhang, Zhexin
Lu, Yida
Wang, Hongning
Qiu, Han
Huang, Minlie
contents Toxicity detection in multimodal text-image content faces growing challenges, especially with multimodal implicit toxicity, where each modality appears benign on its own but conveys hazard when combined. Multimodal implicit toxicity appears not only as formal statements in social platforms but also prompts that can lead to toxic dialogs from Large Vision-Language Models (LVLMs). Despite the success in unimodal text or image moderation, toxicity detection for multimodal content, particularly the multimodal implicit toxicity, remains underexplored. To fill this gap, we comprehensively build a taxonomy for multimodal implicit toxicity (MMIT) and introduce an MMIT-dataset, comprising 2,100 multimodal statements and prompts across 7 risk categories (31 sub-categories) and 5 typical cross-modal correlation modes. To advance the detection of multimodal implicit toxicity, we build ShieldVLM, a model which identifies implicit toxicity in multimodal statements, prompts and dialogs via deliberative cross-modal reasoning. Experiments show that ShieldVLM outperforms existing strong baselines in detecting both implicit and explicit toxicity. The model and dataset will be publicly available to support future researches. Warning: This paper contains potentially sensitive contents.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
Cui, Shiyao
Zhang, Qinglin
Ouyang, Xuan
Chen, Renmiao
Zhang, Zhexin
Lu, Yida
Wang, Hongning
Qiu, Han
Huang, Minlie
Multimedia
Computation and Language
Toxicity detection in multimodal text-image content faces growing challenges, especially with multimodal implicit toxicity, where each modality appears benign on its own but conveys hazard when combined. Multimodal implicit toxicity appears not only as formal statements in social platforms but also prompts that can lead to toxic dialogs from Large Vision-Language Models (LVLMs). Despite the success in unimodal text or image moderation, toxicity detection for multimodal content, particularly the multimodal implicit toxicity, remains underexplored. To fill this gap, we comprehensively build a taxonomy for multimodal implicit toxicity (MMIT) and introduce an MMIT-dataset, comprising 2,100 multimodal statements and prompts across 7 risk categories (31 sub-categories) and 5 typical cross-modal correlation modes. To advance the detection of multimodal implicit toxicity, we build ShieldVLM, a model which identifies implicit toxicity in multimodal statements, prompts and dialogs via deliberative cross-modal reasoning. Experiments show that ShieldVLM outperforms existing strong baselines in detecting both implicit and explicit toxicity. The model and dataset will be publicly available to support future researches. Warning: This paper contains potentially sensitive contents.
title ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
topic Multimedia
Computation and Language
url https://arxiv.org/abs/2505.14035