An Effective, Robust and Fairness-aware Hate Speech Detection Framework

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mou, Guanyi, Lee, Kyumin
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912046551924736
author Mou, Guanyi
Lee, Kyumin
author_facet Mou, Guanyi
Lee, Kyumin
contents With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data insufficiency, estimating model uncertainty, improving robustness against malicious attacks, and handling unintended bias (i.e., fairness). There is an urgent need for accurate, robust, and fair hate speech classification in online social networks. To bridge the gap, we design a data-augmented, fairness addressed, and uncertainty estimated novel framework. As parts of the framework, we propose Bidirectional Quaternion-Quasi-LSTM layers to balance effectiveness and efficiency. To build a generalized model, we combine five datasets collected from three platforms. Experiment results show that our model outperforms eight state-of-the-art methods under both no attack scenario and various attack scenarios, indicating the effectiveness and robustness of our model. We share our code along with combined dataset for better future research
format Preprint
id arxiv_https___arxiv_org_abs_2409_17191
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Effective, Robust and Fairness-aware Hate Speech Detection Framework
Mou, Guanyi
Lee, Kyumin
Computation and Language
Machine Learning
With the widespread online social networks, hate speeches are spreading faster and causing more damage than ever before. Existing hate speech detection methods have limitations in several aspects, such as handling data insufficiency, estimating model uncertainty, improving robustness against malicious attacks, and handling unintended bias (i.e., fairness). There is an urgent need for accurate, robust, and fair hate speech classification in online social networks. To bridge the gap, we design a data-augmented, fairness addressed, and uncertainty estimated novel framework. As parts of the framework, we propose Bidirectional Quaternion-Quasi-LSTM layers to balance effectiveness and efficiency. To build a generalized model, we combine five datasets collected from three platforms. Experiment results show that our model outperforms eight state-of-the-art methods under both no attack scenario and various attack scenarios, indicating the effectiveness and robustness of our model. We share our code along with combined dataset for better future research
title An Effective, Robust and Fairness-aware Hate Speech Detection Framework
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2409.17191