Reasoning-Aware Multimodal Fusion for Hateful Video Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Shuonan, Chen, Tailin, Yue, Jiangbei, Cheng, Guangliang, Jiao, Jianbo, Fu, Zeyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914614697000960
author Yang, Shuonan
Chen, Tailin
Yue, Jiangbei
Cheng, Guangliang
Jiao, Jianbo
Fu, Zeyu
author_facet Yang, Shuonan
Chen, Tailin
Yue, Jiangbei
Cheng, Guangliang
Jiao, Jianbo
Fu, Zeyu
contents Hate speech in online videos is posing an increasingly serious threat to digital platforms, especially as video content becomes increasingly multimodal and context-dependent. Existing methods often struggle to effectively fuse the complex semantic relationships between modalities and lack the ability to understand nuanced hateful content. To address these issues, we propose an innovative Reasoning-Aware Multimodal Fusion (RAMF) framework. To tackle the first challenge, we design Local-Global Context Fusion (LGCF) to capture both local salient cues and global temporal structures, and propose Semantic Cross Attention (SCA) to enable fine-grained multimodal semantic interaction. To tackle the second challenge, we introduce adversarial reasoning-a structured three-stage process where a vision-language model generates (i) objective descriptions, (ii) hate-assumed inferences, and (iii) non-hate-assumed inferences-providing complementary semantic perspectives that enrich the model's contextual understanding of nuanced hateful intent. Evaluations on two real-world hateful video datasets demonstrate that our method achieves robust generalisation performance, improving upon state-of-the-art methods by 3% and 7% in Macro-F1 and hate class recall, respectively. The source codes and data required to reproduce our results are available at https://github.com/Multimodal-Intelligence-Lab-MIL/RAMF.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02743
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reasoning-Aware Multimodal Fusion for Hateful Video Detection
Yang, Shuonan
Chen, Tailin
Yue, Jiangbei
Cheng, Guangliang
Jiao, Jianbo
Fu, Zeyu
Computer Vision and Pattern Recognition
Artificial Intelligence
Hate speech in online videos is posing an increasingly serious threat to digital platforms, especially as video content becomes increasingly multimodal and context-dependent. Existing methods often struggle to effectively fuse the complex semantic relationships between modalities and lack the ability to understand nuanced hateful content. To address these issues, we propose an innovative Reasoning-Aware Multimodal Fusion (RAMF) framework. To tackle the first challenge, we design Local-Global Context Fusion (LGCF) to capture both local salient cues and global temporal structures, and propose Semantic Cross Attention (SCA) to enable fine-grained multimodal semantic interaction. To tackle the second challenge, we introduce adversarial reasoning-a structured three-stage process where a vision-language model generates (i) objective descriptions, (ii) hate-assumed inferences, and (iii) non-hate-assumed inferences-providing complementary semantic perspectives that enrich the model's contextual understanding of nuanced hateful intent. Evaluations on two real-world hateful video datasets demonstrate that our method achieves robust generalisation performance, improving upon state-of-the-art methods by 3% and 7% in Macro-F1 and hate class recall, respectively. The source codes and data required to reproduce our results are available at https://github.com/Multimodal-Intelligence-Lab-MIL/RAMF.
title Reasoning-Aware Multimodal Fusion for Hateful Video Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2512.02743