Hate Speech Detection Using Cross-Platform Social Media Data In English and German Language

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shahi, Gautam Kishore, Majchrzak, Tim A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914966962962432
author Shahi, Gautam Kishore
Majchrzak, Tim A.
author_facet Shahi, Gautam Kishore
Majchrzak, Tim A.
contents Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is yet unaccomplished. The challenge for hate speech detection as text classification is the cost of obtaining high-quality training data. This study focuses on detecting bilingual hate speech in YouTube comments and measuring the impact of using additional data from other platforms in the performance of the classification model. We examine the value of additional training datasets from cross-platforms for improving the performance of classification models. We also included factors such as content similarity, definition similarity, and common hate words to measure the impact of datasets on performance. Our findings show that adding more similar datasets based on content similarity, hate words, and definitions improves the performance of classification models. The best performance was obtained by combining datasets from YouTube comments, Twitter, and Gab with an F1-score of 0.74 and 0.68 for English and German YouTube comments.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05287
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hate Speech Detection Using Cross-Platform Social Media Data In English and German Language
Shahi, Gautam Kishore
Majchrzak, Tim A.
Computation and Language
Artificial Intelligence
Social and Information Networks
Hate speech has grown into a pervasive phenomenon, intensifying during times of crisis, elections, and social unrest. Multiple approaches have been developed to detect hate speech using artificial intelligence, but a generalized model is yet unaccomplished. The challenge for hate speech detection as text classification is the cost of obtaining high-quality training data. This study focuses on detecting bilingual hate speech in YouTube comments and measuring the impact of using additional data from other platforms in the performance of the classification model. We examine the value of additional training datasets from cross-platforms for improving the performance of classification models. We also included factors such as content similarity, definition similarity, and common hate words to measure the impact of datasets on performance. Our findings show that adding more similar datasets based on content similarity, hate words, and definitions improves the performance of classification models. The best performance was obtained by combining datasets from YouTube comments, Twitter, and Gab with an F1-score of 0.74 and 0.68 for English and German YouTube comments.
title Hate Speech Detection Using Cross-Platform Social Media Data In English and German Language
topic Computation and Language
Artificial Intelligence
Social and Information Networks
url https://arxiv.org/abs/2410.05287