Hate Speech Detection and Classification in Amharic Text with Deep Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gashe, Samuel Minale, Yimam, Seid Muhie, Assabie, Yaregal
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911980872269824
author Gashe, Samuel Minale
Yimam, Seid Muhie
Assabie, Yaregal
author_facet Gashe, Samuel Minale
Yimam, Seid Muhie
Assabie, Yaregal
contents Hate speech is a growing problem on social media. It can seriously impact society, especially in countries like Ethiopia, where it can trigger conflicts among diverse ethnic and religious groups. While hate speech detection in resource rich languages are progressing, for low resource languages such as Amharic are lacking. To address this gap, we develop Amharic hate speech data and SBi-LSTM deep learning model that can detect and classify text into four categories of hate speech: racial, religious, gender, and non-hate speech. We have annotated 5k Amharic social media post and comment data into four categories. The data is annotated using a custom annotation tool by a total of 100 native Amharic speakers. The model achieves a 94.8 F1-score performance. Future improvements will include expanding the dataset and develop state-of-the art models. Keywords: Amharic hate speech detection, classification, Amharic dataset, Deep Learning, SBi-LSTM
format Preprint
id arxiv_https___arxiv_org_abs_2408_03849
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hate Speech Detection and Classification in Amharic Text with Deep Learning
Gashe, Samuel Minale
Yimam, Seid Muhie
Assabie, Yaregal
Computation and Language
Machine Learning
Hate speech is a growing problem on social media. It can seriously impact society, especially in countries like Ethiopia, where it can trigger conflicts among diverse ethnic and religious groups. While hate speech detection in resource rich languages are progressing, for low resource languages such as Amharic are lacking. To address this gap, we develop Amharic hate speech data and SBi-LSTM deep learning model that can detect and classify text into four categories of hate speech: racial, religious, gender, and non-hate speech. We have annotated 5k Amharic social media post and comment data into four categories. The data is annotated using a custom annotation tool by a total of 100 native Amharic speakers. The model achieves a 94.8 F1-score performance. Future improvements will include expanding the dataset and develop state-of-the art models. Keywords: Amharic hate speech detection, classification, Amharic dataset, Deep Learning, SBi-LSTM
title Hate Speech Detection and Classification in Amharic Text with Deep Learning
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2408.03849