Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Yinghui, Chen, Tailin, Zhang, Yuchen, Fu, Zeyu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909614777303040
author Zhang, Yinghui
Chen, Tailin
Zhang, Yuchen
Fu, Zeyu
author_facet Zhang, Yinghui
Chen, Tailin
Zhang, Yuchen
Fu, Zeyu
contents The rapid rise of video content on platforms such as TikTok and YouTube has transformed information dissemination, but it has also facilitated the spread of harmful content, particularly hate videos. Despite significant efforts to combat hate speech, detecting these videos remains challenging due to their often implicit nature. Current detection methods primarily rely on unimodal approaches, which inadequately capture the complementary features across different modalities. While multimodal techniques offer a broader perspective, many fail to effectively integrate temporal dynamics and modality-wise interactions essential for identifying nuanced hate content. In this paper, we present CMFusion, an enhanced multimodal hate video detection model utilizing a novel Channel-wise and Modality-wise Fusion Mechanism. CMFusion first extracts features from text, audio, and video modalities using pre-trained models and then incorporates a temporal cross-attention mechanism to capture dependencies between video and audio streams. The learned features are then processed by channel-wise and modality-wise fusion modules to obtain informative representations of videos. Our extensive experiments on a real-world dataset demonstrate that CMFusion significantly outperforms five widely used baselines in terms of accuracy, precision, recall, and F1 score. Comprehensive ablation studies and parameter analyses further validate our design choices, highlighting the model's effectiveness in detecting hate videos. The source codes will be made publicly available at https://github.com/EvelynZ10/cmfusion.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12051
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
Zhang, Yinghui
Chen, Tailin
Zhang, Yuchen
Fu, Zeyu
Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
The rapid rise of video content on platforms such as TikTok and YouTube has transformed information dissemination, but it has also facilitated the spread of harmful content, particularly hate videos. Despite significant efforts to combat hate speech, detecting these videos remains challenging due to their often implicit nature. Current detection methods primarily rely on unimodal approaches, which inadequately capture the complementary features across different modalities. While multimodal techniques offer a broader perspective, many fail to effectively integrate temporal dynamics and modality-wise interactions essential for identifying nuanced hate content. In this paper, we present CMFusion, an enhanced multimodal hate video detection model utilizing a novel Channel-wise and Modality-wise Fusion Mechanism. CMFusion first extracts features from text, audio, and video modalities using pre-trained models and then incorporates a temporal cross-attention mechanism to capture dependencies between video and audio streams. The learned features are then processed by channel-wise and modality-wise fusion modules to obtain informative representations of videos. Our extensive experiments on a real-world dataset demonstrate that CMFusion significantly outperforms five widely used baselines in terms of accuracy, precision, recall, and F1 score. Comprehensive ablation studies and parameter analyses further validate our design choices, highlighting the model's effectiveness in detecting hate videos. The source codes will be made publicly available at https://github.com/EvelynZ10/cmfusion.
title Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
topic Multimedia
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.12051