Saved in:
Bibliographic Details
Main Authors: Jin, Bohan, Qi, Shuhan, Chen, Kehai, Guo, Xinyi, Wang, Xuan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.17144
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909620849606656
author Jin, Bohan
Qi, Shuhan
Chen, Kehai
Guo, Xinyi
Wang, Xuan
author_facet Jin, Bohan
Qi, Shuhan
Chen, Kehai
Guo, Xinyi
Wang, Xuan
contents The widespread use of Large Multimodal Models (LMMs) has raised concerns about model toxicity. However, current research mainly focuses on explicit toxicity, with less attention to some more implicit toxicity regarding prejudice and discrimination. To address this limitation, we introduce a subtler type of toxicity named dual-implicit toxicity and a novel toxicity benchmark termed MDIT-Bench: Multimodal Dual-Implicit Toxicity Benchmark. Specifically, we first create the MDIT-Dataset with dual-implicit toxicity using the proposed Multi-stage Human-in-loop In-context Generation method. Based on this dataset, we construct the MDIT-Bench, a benchmark for evaluating the sensitivity of models to dual-implicit toxicity, with 317,638 questions covering 12 categories, 23 subcategories, and 780 topics. MDIT-Bench includes three difficulty levels, and we propose a metric to measure the toxicity gap exhibited by the model across them. In the experiment, we conducted MDIT-Bench on 13 prominent LMMs, and the results show that these LMMs cannot handle dual-implicit toxicity effectively. The model's performance drops significantly in hard level, revealing that these LMMs still contain a significant amount of hidden but activatable toxicity. Data are available at https://github.com/nuo1nuo/MDIT-Bench.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17144
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
Jin, Bohan
Qi, Shuhan
Chen, Kehai
Guo, Xinyi
Wang, Xuan
Computation and Language
Artificial Intelligence
The widespread use of Large Multimodal Models (LMMs) has raised concerns about model toxicity. However, current research mainly focuses on explicit toxicity, with less attention to some more implicit toxicity regarding prejudice and discrimination. To address this limitation, we introduce a subtler type of toxicity named dual-implicit toxicity and a novel toxicity benchmark termed MDIT-Bench: Multimodal Dual-Implicit Toxicity Benchmark. Specifically, we first create the MDIT-Dataset with dual-implicit toxicity using the proposed Multi-stage Human-in-loop In-context Generation method. Based on this dataset, we construct the MDIT-Bench, a benchmark for evaluating the sensitivity of models to dual-implicit toxicity, with 317,638 questions covering 12 categories, 23 subcategories, and 780 topics. MDIT-Bench includes three difficulty levels, and we propose a metric to measure the toxicity gap exhibited by the model across them. In the experiment, we conducted MDIT-Bench on 13 prominent LMMs, and the results show that these LMMs cannot handle dual-implicit toxicity effectively. The model's performance drops significantly in hard level, revealing that these LMMs still contain a significant amount of hidden but activatable toxicity. Data are available at https://github.com/nuo1nuo/MDIT-Bench.
title MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.17144