Analysis and Detection of Multilingual Hate Speech Using Transformer Based Deep Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Das, Arijit, Nandy, Somashree, Saha, Rupam, Das, Srijan, Saha, Diganta
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916162797830144
author Das, Arijit
Nandy, Somashree
Saha, Rupam
Das, Srijan
Saha, Diganta
author_facet Das, Arijit
Nandy, Somashree
Saha, Rupam
Das, Srijan
Saha, Diganta
contents Hate speech is harmful content that directly attacks or promotes hatred against members of groups or individuals based on actual or perceived aspects of identity, such as racism, religion, or sexual orientation. This can affect social life on social media platforms as hateful content shared through social media can harm both individuals and communities. As the prevalence of hate speech increases online, the demand for automated detection as an NLP task is increasing. In this work, the proposed method is using transformer-based model to detect hate speech in social media, like twitter, Facebook, WhatsApp, Instagram, etc. The proposed model is independent of languages and has been tested on Italian, English, German, Bengali. The Gold standard datasets were collected from renowned researcher Zeerak Talat, Sara Tonelli, Melanie Siegel, and Rezaul Karim. The success rate of the proposed model for hate speech detection is higher than the existing baseline and state-of-the-art models with accuracy in Bengali dataset is 89%, in English: 91%, in German dataset 91% and in Italian dataset it is 77%. The proposed algorithm shows substantial improvement to the benchmark method.
format Preprint
id arxiv_https___arxiv_org_abs_2401_11021
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Analysis and Detection of Multilingual Hate Speech Using Transformer Based Deep Learning
Das, Arijit
Nandy, Somashree
Saha, Rupam
Das, Srijan
Saha, Diganta
Computation and Language
Artificial Intelligence
Information Retrieval
Hate speech is harmful content that directly attacks or promotes hatred against members of groups or individuals based on actual or perceived aspects of identity, such as racism, religion, or sexual orientation. This can affect social life on social media platforms as hateful content shared through social media can harm both individuals and communities. As the prevalence of hate speech increases online, the demand for automated detection as an NLP task is increasing. In this work, the proposed method is using transformer-based model to detect hate speech in social media, like twitter, Facebook, WhatsApp, Instagram, etc. The proposed model is independent of languages and has been tested on Italian, English, German, Bengali. The Gold standard datasets were collected from renowned researcher Zeerak Talat, Sara Tonelli, Melanie Siegel, and Rezaul Karim. The success rate of the proposed model for hate speech detection is higher than the existing baseline and state-of-the-art models with accuracy in Bengali dataset is 89%, in English: 91%, in German dataset 91% and in Italian dataset it is 77%. The proposed algorithm shows substantial improvement to the benchmark method.
title Analysis and Detection of Multilingual Hate Speech Using Transformer Based Deep Learning
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2401.11021