ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916866664955904 |
|---|---|
| author | Kashyap, Gautam Siddharth Azeez, Mohammad Anas Ali, Rafiq Siddiqui, Zohaib Hasan Gao, Jiechao Naseem, Usman |
| author_facet | Kashyap, Gautam Siddharth Azeez, Mohammad Anas Ali, Rafiq Siddiqui, Zohaib Hasan Gao, Jiechao Naseem, Usman |
| contents | Hate speech targeting children on social media is a serious and growing problem, yet current NLP systems struggle to detect it effectively. This gap exists mainly because existing datasets focus on adults, lack age specific labels, miss nuanced linguistic cues, and are often too small for robust modeling. To address this, we introduce ChildGuard, the first large scale English dataset dedicated to hate speech aimed at children. It contains 351,877 annotated examples from X (formerly Twitter), Reddit, and YouTube, labeled by three age groups: younger children (under 11), pre teens (11--12), and teens (13--17). The dataset is split into two subsets for fine grained analysis: a contextual subset (157K) focusing on discourse level features, and a lexical subset (194K) emphasizing word-level sentiment and vocabulary. Benchmarking state of the art hate speech models on ChildGuard reveals notable drops in performance, highlighting the challenges of detecting child directed hate speech. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21613 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech Kashyap, Gautam Siddharth Azeez, Mohammad Anas Ali, Rafiq Siddiqui, Zohaib Hasan Gao, Jiechao Naseem, Usman Computation and Language Sound Audio and Speech Processing Hate speech targeting children on social media is a serious and growing problem, yet current NLP systems struggle to detect it effectively. This gap exists mainly because existing datasets focus on adults, lack age specific labels, miss nuanced linguistic cues, and are often too small for robust modeling. To address this, we introduce ChildGuard, the first large scale English dataset dedicated to hate speech aimed at children. It contains 351,877 annotated examples from X (formerly Twitter), Reddit, and YouTube, labeled by three age groups: younger children (under 11), pre teens (11--12), and teens (13--17). The dataset is split into two subsets for fine grained analysis: a contextual subset (157K) focusing on discourse level features, and a lexical subset (194K) emphasizing word-level sentiment and vocabulary. Benchmarking state of the art hate speech models on ChildGuard reveals notable drops in performance, highlighting the challenges of detecting child directed hate speech. |
| title | ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.21613 |