HateDebias: On the Diversity and Variability of Hate Speech Debiasing

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Hongyan, Chen, Zhengming, Li, Zijian, Lin, Nankai, Wang, Lianxi, Jiang, Shengyi, Yang, Aimin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908502581051392
author Wu, Hongyan
Chen, Zhengming
Li, Zijian
Lin, Nankai
Wang, Lianxi
Jiang, Shengyi
Yang, Aimin
author_facet Wu, Hongyan
Chen, Zhengming
Li, Zijian
Lin, Nankai
Wang, Lianxi
Jiang, Shengyi
Yang, Aimin
contents Hate speech frequently appears on social media platforms and urgently needs to be effectively controlled. Alleviating the bias caused by hate speech can help resolve various ethical issues. Although existing research has constructed several datasets for hate speech detection, these datasets seldom consider the diversity and variability of bias, making them far from real-world scenarios. To fill this gap, we propose a benchmark HateDebias to analyze the fairness of models under dynamically evolving environments. Specifically, to meet the diversity of biases, we collect hate speech data with different types of biases from real-world scenarios. To further simulate the variability in the real-world scenarios(i.e., the changing of bias attributes in datasets), we construct a dataset to follow the continuous learning setting and evaluate the detection accuracy of models on the HateDebias, where performance degradation indicates a significant bias toward a specific attribute. To provide a potential direction, we further propose a continual debiasing framework tailored to dynamic bias in real-world scenarios, integrating memory replay and bias information regularization to ensure the fairness of the model. Experiment results on the HateDebias benchmark reveal that our methods achieve improved performance in mitigating dynamic biases in real-world scenarios, highlighting the practicality in real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2406_04876
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle HateDebias: On the Diversity and Variability of Hate Speech Debiasing
Wu, Hongyan
Chen, Zhengming
Li, Zijian
Lin, Nankai
Wang, Lianxi
Jiang, Shengyi
Yang, Aimin
Computation and Language
Hate speech frequently appears on social media platforms and urgently needs to be effectively controlled. Alleviating the bias caused by hate speech can help resolve various ethical issues. Although existing research has constructed several datasets for hate speech detection, these datasets seldom consider the diversity and variability of bias, making them far from real-world scenarios. To fill this gap, we propose a benchmark HateDebias to analyze the fairness of models under dynamically evolving environments. Specifically, to meet the diversity of biases, we collect hate speech data with different types of biases from real-world scenarios. To further simulate the variability in the real-world scenarios(i.e., the changing of bias attributes in datasets), we construct a dataset to follow the continuous learning setting and evaluate the detection accuracy of models on the HateDebias, where performance degradation indicates a significant bias toward a specific attribute. To provide a potential direction, we further propose a continual debiasing framework tailored to dynamic bias in real-world scenarios, integrating memory replay and bias information regularization to ensure the fairness of the model. Experiment results on the HateDebias benchmark reveal that our methods achieve improved performance in mitigating dynamic biases in real-world scenarios, highlighting the practicality in real-world applications.
title HateDebias: On the Diversity and Variability of Hate Speech Debiasing
topic Computation and Language
url https://arxiv.org/abs/2406.04876