ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kashyap, Gautam Siddharth, Azeez, Mohammad Anas, Ali, Rafiq, Siddiqui, Zohaib Hasan, Gao, Jiechao, Naseem, Usman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916866664955904
author Kashyap, Gautam Siddharth
Azeez, Mohammad Anas
Ali, Rafiq
Siddiqui, Zohaib Hasan
Gao, Jiechao
Naseem, Usman
author_facet Kashyap, Gautam Siddharth
Azeez, Mohammad Anas
Ali, Rafiq
Siddiqui, Zohaib Hasan
Gao, Jiechao
Naseem, Usman
contents Hate speech targeting children on social media is a serious and growing problem, yet current NLP systems struggle to detect it effectively. This gap exists mainly because existing datasets focus on adults, lack age specific labels, miss nuanced linguistic cues, and are often too small for robust modeling. To address this, we introduce ChildGuard, the first large scale English dataset dedicated to hate speech aimed at children. It contains 351,877 annotated examples from X (formerly Twitter), Reddit, and YouTube, labeled by three age groups: younger children (under 11), pre teens (11--12), and teens (13--17). The dataset is split into two subsets for fine grained analysis: a contextual subset (157K) focusing on discourse level features, and a lexical subset (194K) emphasizing word-level sentiment and vocabulary. Benchmarking state of the art hate speech models on ChildGuard reveals notable drops in performance, highlighting the challenges of detecting child directed hate speech.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21613
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
Kashyap, Gautam Siddharth
Azeez, Mohammad Anas
Ali, Rafiq
Siddiqui, Zohaib Hasan
Gao, Jiechao
Naseem, Usman
Computation and Language
Sound
Audio and Speech Processing
Hate speech targeting children on social media is a serious and growing problem, yet current NLP systems struggle to detect it effectively. This gap exists mainly because existing datasets focus on adults, lack age specific labels, miss nuanced linguistic cues, and are often too small for robust modeling. To address this, we introduce ChildGuard, the first large scale English dataset dedicated to hate speech aimed at children. It contains 351,877 annotated examples from X (formerly Twitter), Reddit, and YouTube, labeled by three age groups: younger children (under 11), pre teens (11--12), and teens (13--17). The dataset is split into two subsets for fine grained analysis: a contextual subset (157K) focusing on discourse level features, and a lexical subset (194K) emphasizing word-level sentiment and vocabulary. Benchmarking state of the art hate speech models on ChildGuard reveals notable drops in performance, highlighting the challenges of detecting child directed hate speech.
title ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2506.21613