Automatic Textual Normalization for Hate Speech Detection
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911965890215936 |
|---|---|
| author | Nguyen, Anh Thi-Hoang Nguyen, Dung Ha Nguyen, Nguyet Thi Ho, Khanh Thanh-Duy Van Nguyen, Kiet |
| author_facet | Nguyen, Anh Thi-Hoang Nguyen, Dung Ha Nguyen, Nguyet Thi Ho, Khanh Thanh-Duy Van Nguyen, Kiet |
| contents | Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the Vietnamese language address this issue as a problem of lexical normalization, involving the creation of manual rules or the implementation of multi-staged deep learning frameworks, which necessitate extensive efforts to craft intricate rules. In contrast, our approach is straightforward, employing solely a sequence-to-sequence (Seq2Seq) model. In this research, we provide a dataset for textual normalization, comprising 2,181 human-annotated comments with an inter-annotator agreement of 0.9014. By leveraging the Seq2Seq model for textual normalization, our results reveal that the accuracy achieved falls slightly short of 70%. Nevertheless, textual normalization enhances the accuracy of the Hate Speech Detection (HSD) task by approximately 2%, demonstrating its potential to improve the performance of complex NLP tasks. Our dataset is accessible for research purposes. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_06851 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Automatic Textual Normalization for Hate Speech Detection Nguyen, Anh Thi-Hoang Nguyen, Dung Ha Nguyen, Nguyet Thi Ho, Khanh Thanh-Duy Van Nguyen, Kiet Computation and Language Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the Vietnamese language address this issue as a problem of lexical normalization, involving the creation of manual rules or the implementation of multi-staged deep learning frameworks, which necessitate extensive efforts to craft intricate rules. In contrast, our approach is straightforward, employing solely a sequence-to-sequence (Seq2Seq) model. In this research, we provide a dataset for textual normalization, comprising 2,181 human-annotated comments with an inter-annotator agreement of 0.9014. By leveraging the Seq2Seq model for textual normalization, our results reveal that the accuracy achieved falls slightly short of 70%. Nevertheless, textual normalization enhances the accuracy of the Hate Speech Detection (HSD) task by approximately 2%, demonstrating its potential to improve the performance of complex NLP tasks. Our dataset is accessible for research purposes. |
| title | Automatic Textual Normalization for Hate Speech Detection |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2311.06851 |