DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dalal, Dwip, Vashishtha, Gautam, Rani, Anku, Reganti, Aishwarya, Patwa, Parth, Sarique, Mohd, Gupta, Chandan, Nath, Keshav, Reddy, Viswanatha, Jain, Vinija, Chadha, Aman, Das, Amitava, Sheth, Amit, Ekbal, Asif
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911200967655424
author Dalal, Dwip
Vashishtha, Gautam
Rani, Anku
Reganti, Aishwarya
Patwa, Parth
Sarique, Mohd
Gupta, Chandan
Nath, Keshav
Reddy, Viswanatha
Jain, Vinija
Chadha, Aman
Das, Amitava
Sheth, Amit
Ekbal, Asif
author_facet Dalal, Dwip
Vashishtha, Gautam
Rani, Anku
Reganti, Aishwarya
Patwa, Parth
Sarique, Mohd
Gupta, Chandan
Nath, Keshav
Reddy, Viswanatha
Jain, Vinija
Chadha, Aman
Das, Amitava
Sheth, Amit
Ekbal, Asif
contents The rise in harmful online content not only distorts public discourse but also poses significant challenges to maintaining a healthy digital environment. In response to this, we introduce a multimodal dataset uniquely crafted for identifying hate in digital content. Central to our methodology is the innovative application of watermarked, stability-enhanced, stable diffusion techniques combined with the Digital Attention Analysis Module (DAAM). This combination is instrumental in pinpointing the hateful elements within images, thereby generating detailed hate attention maps, which are used to blur these regions from the image, thereby removing the hateful sections of the image. We release this data set as a part of the dehate shared task. This paper also describes the details of the shared task. Furthermore, we present DeHater, a vision-language model designed for multimodal dehatification tasks. Our approach sets a new standard in AI-driven image hate detection given textual prompts, contributing to the development of more ethical AI applications in social media.
format Preprint
id arxiv_https___arxiv_org_abs_2509_21787
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
Dalal, Dwip
Vashishtha, Gautam
Rani, Anku
Reganti, Aishwarya
Patwa, Parth
Sarique, Mohd
Gupta, Chandan
Nath, Keshav
Reddy, Viswanatha
Jain, Vinija
Chadha, Aman
Das, Amitava
Sheth, Amit
Ekbal, Asif
Computer Vision and Pattern Recognition
Computation and Language
The rise in harmful online content not only distorts public discourse but also poses significant challenges to maintaining a healthy digital environment. In response to this, we introduce a multimodal dataset uniquely crafted for identifying hate in digital content. Central to our methodology is the innovative application of watermarked, stability-enhanced, stable diffusion techniques combined with the Digital Attention Analysis Module (DAAM). This combination is instrumental in pinpointing the hateful elements within images, thereby generating detailed hate attention maps, which are used to blur these regions from the image, thereby removing the hateful sections of the image. We release this data set as a part of the dehate shared task. This paper also describes the details of the shared task. Furthermore, we present DeHater, a vision-language model designed for multimodal dehatification tasks. Our approach sets a new standard in AI-driven image hate detection given textual prompts, contributing to the development of more ethical AI applications in social media.
title DeHate: A Stable Diffusion-based Multimodal Approach to Mitigate Hate Speech in Images
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2509.21787