HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915447238033408 |
|---|---|
| author | Wang, Han Wang, Zhuoran Lee, Roy Ka-Wei |
| author_facet | Wang, Han Wang, Zhuoran Lee, Roy Ka-Wei |
| contents | Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level and segment-level annotations, comprising over 11,714 segments labeled as Normal or across five Offensive categories: Hateful, Insulting, Sexual, Violence, Self-Harm, along with explicit target victim labels. Our three-stage annotation process yields high inter-annotator agreement (Krippendorff's alpha = 0.817). We propose three tasks to benchmark performance: (1) Trimmed Hateful Video Classification, (2) Temporal Hateful Video Localization, and (3) Online Hateful Video Classification. Results highlight substantial gaps in current models, emphasizing the need for more sophisticated multimodal and temporally aware approaches. The HateClipSeg dataset are publicly available at https://github.com/Social-AI-Studio/HateClipSeg.git. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_01712 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection Wang, Han Wang, Zhuoran Lee, Roy Ka-Wei Computer Vision and Pattern Recognition Artificial Intelligence cs.CV, cs.MM I.2.10 Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level and segment-level annotations, comprising over 11,714 segments labeled as Normal or across five Offensive categories: Hateful, Insulting, Sexual, Violence, Self-Harm, along with explicit target victim labels. Our three-stage annotation process yields high inter-annotator agreement (Krippendorff's alpha = 0.817). We propose three tasks to benchmark performance: (1) Trimmed Hateful Video Classification, (2) Temporal Hateful Video Localization, and (3) Online Hateful Video Classification. Results highlight substantial gaps in current models, emphasizing the need for more sophisticated multimodal and temporally aware approaches. The HateClipSeg dataset are publicly available at https://github.com/Social-AI-Studio/HateClipSeg.git. |
| title | HateClipSeg: A Segment-Level Annotated Dataset for Fine-Grained Hate Video Detection |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence cs.CV, cs.MM I.2.10 |
| url | https://arxiv.org/abs/2508.01712 |