Automatic Textual Normalization for Hate Speech Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Anh Thi-Hoang, Nguyen, Dung Ha, Nguyen, Nguyet Thi, Ho, Khanh Thanh-Duy, Van Nguyen, Kiet
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911965890215936
author Nguyen, Anh Thi-Hoang
Nguyen, Dung Ha
Nguyen, Nguyet Thi
Ho, Khanh Thanh-Duy
Van Nguyen, Kiet
author_facet Nguyen, Anh Thi-Hoang
Nguyen, Dung Ha
Nguyen, Nguyet Thi
Ho, Khanh Thanh-Duy
Van Nguyen, Kiet
contents Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the Vietnamese language address this issue as a problem of lexical normalization, involving the creation of manual rules or the implementation of multi-staged deep learning frameworks, which necessitate extensive efforts to craft intricate rules. In contrast, our approach is straightforward, employing solely a sequence-to-sequence (Seq2Seq) model. In this research, we provide a dataset for textual normalization, comprising 2,181 human-annotated comments with an inter-annotator agreement of 0.9014. By leveraging the Seq2Seq model for textual normalization, our results reveal that the accuracy achieved falls slightly short of 70%. Nevertheless, textual normalization enhances the accuracy of the Hate Speech Detection (HSD) task by approximately 2%, demonstrating its potential to improve the performance of complex NLP tasks. Our dataset is accessible for research purposes.
format Preprint
id arxiv_https___arxiv_org_abs_2311_06851
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Automatic Textual Normalization for Hate Speech Detection
Nguyen, Anh Thi-Hoang
Nguyen, Dung Ha
Nguyen, Nguyet Thi
Ho, Khanh Thanh-Duy
Van Nguyen, Kiet
Computation and Language
Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the Vietnamese language address this issue as a problem of lexical normalization, involving the creation of manual rules or the implementation of multi-staged deep learning frameworks, which necessitate extensive efforts to craft intricate rules. In contrast, our approach is straightforward, employing solely a sequence-to-sequence (Seq2Seq) model. In this research, we provide a dataset for textual normalization, comprising 2,181 human-annotated comments with an inter-annotator agreement of 0.9014. By leveraging the Seq2Seq model for textual normalization, our results reveal that the accuracy achieved falls slightly short of 70%. Nevertheless, textual normalization enhances the accuracy of the Hate Speech Detection (HSD) task by approximately 2%, demonstrating its potential to improve the performance of complex NLP tasks. Our dataset is accessible for research purposes.
title Automatic Textual Normalization for Hate Speech Detection
topic Computation and Language
url https://arxiv.org/abs/2311.06851