MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Piot, Paloma, Martín-Rodilla, Patricia, Parapar, Javier
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915271658176512
author Piot, Paloma
Martín-Rodilla, Patricia
Parapar, Javier
author_facet Piot, Paloma
Martín-Rodilla, Patricia
Parapar, Javier
contents Hate speech represents a pervasive and detrimental form of online discourse, often manifested through an array of slurs, from hateful tweets to defamatory posts. As such speech proliferates, it connects people globally and poses significant social, psychological, and occasionally physical threats to targeted individuals and communities. Current computational linguistic approaches for tackling this phenomenon rely on labelled social media datasets for training. For unifying efforts, our study advances in the critical need for a comprehensive meta-collection, advocating for an extensive dataset to help counteract this problem effectively. We scrutinized over 60 datasets, selectively integrating those pertinent into MetaHate. This paper offers a detailed examination of existing collections, highlighting their strengths and limitations. Our findings contribute to a deeper understanding of the existing datasets, paving the way for training more robust and adaptable models. These enhanced models are essential for effectively combating the dynamic and complex nature of hate speech in the digital realm.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06526
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
Piot, Paloma
Martín-Rodilla, Patricia
Parapar, Javier
Computation and Language
Social and Information Networks
Hate speech represents a pervasive and detrimental form of online discourse, often manifested through an array of slurs, from hateful tweets to defamatory posts. As such speech proliferates, it connects people globally and poses significant social, psychological, and occasionally physical threats to targeted individuals and communities. Current computational linguistic approaches for tackling this phenomenon rely on labelled social media datasets for training. For unifying efforts, our study advances in the critical need for a comprehensive meta-collection, advocating for an extensive dataset to help counteract this problem effectively. We scrutinized over 60 datasets, selectively integrating those pertinent into MetaHate. This paper offers a detailed examination of existing collections, highlighting their strengths and limitations. Our findings contribute to a deeper understanding of the existing datasets, paving the way for training more robust and adaptable models. These enhanced models are essential for effectively combating the dynamic and complex nature of hate speech in the digital realm.
title MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
topic Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2401.06526