HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Han, Wang, Zhuoran, Lee, Roy Ka-Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915447238033408
author Wang, Han
Wang, Zhuoran
Lee, Roy Ka-Wei
author_facet Wang, Han
Wang, Zhuoran
Lee, Roy Ka-Wei
contents Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level and segment-level annotations, comprising over 11,714 segments labeled as Normal or across five Offensive categories: Hateful, Insulting, Sexual, Violence, Self-Harm, along with explicit target victim labels. Our three-stage annotation process yields high inter-annotator agreement (Krippendorff's alpha = 0.817). We propose three tasks to benchmark performance: (1) Trimmed Hateful Video Classification, (2) Temporal Hateful Video Localization, and (3) Online Hateful Video Classification. Results highlight substantial gaps in current models, emphasizing the need for more sophisticated multimodal and temporally aware approaches. The HateClipSeg dataset are publicly available at https://github.com/Social-AI-Studio/HateClipSeg.git.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection
Wang, Han
Wang, Zhuoran
Lee, Roy Ka-Wei
Computer Vision and Pattern Recognition
Artificial Intelligence
cs.CV, cs.MM
I.2.10
Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level and segment-level annotations, comprising over 11,714 segments labeled as Normal or across five Offensive categories: Hateful, Insulting, Sexual, Violence, Self-Harm, along with explicit target victim labels. Our three-stage annotation process yields high inter-annotator agreement (Krippendorff's alpha = 0.817). We propose three tasks to benchmark performance: (1) Trimmed Hateful Video Classification, (2) Temporal Hateful Video Localization, and (3) Online Hateful Video Classification. Results highlight substantial gaps in current models, emphasizing the need for more sophisticated multimodal and temporally aware approaches. The HateClipSeg dataset are publicly available at https://github.com/Social-AI-Studio/HateClipSeg.git.
title HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
cs.CV, cs.MM
I.2.10
url https://arxiv.org/abs/2508.01712