Evaluating Simple Debiasing Techniques in RoBERTa-based Hate Speech Detection Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Iftimie, Diana, Zinn, Erik
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916584475328512
author Iftimie, Diana
Zinn, Erik
author_facet Iftimie, Diana
Zinn, Erik
contents The hate speech detection task is known to suffer from bias against African American English (AAE) dialect text, due to the annotation bias present in the underlying hate speech datasets used to train these models. This leads to a disparity where normal AAE text is more likely to be misclassified as abusive/hateful compared to non-AAE text. Simple debiasing techniques have been developed in the past to counter this sort of disparity, and in this work, we apply and evaluate these techniques in the scope of RoBERTa-based encoders. Experimental results suggest that the success of these techniques depends heavily on the methods used for training dataset construction, but with proper consideration of representation bias, they can reduce the disparity seen among dialect subgroups on the hate speech detection task.
format Preprint
id arxiv_https___arxiv_org_abs_2501_15430
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating Simple Debiasing Techniques in RoBERTa-based Hate Speech Detection Models
Iftimie, Diana
Zinn, Erik
Computation and Language
The hate speech detection task is known to suffer from bias against African American English (AAE) dialect text, due to the annotation bias present in the underlying hate speech datasets used to train these models. This leads to a disparity where normal AAE text is more likely to be misclassified as abusive/hateful compared to non-AAE text. Simple debiasing techniques have been developed in the past to counter this sort of disparity, and in this work, we apply and evaluate these techniques in the scope of RoBERTa-based encoders. Experimental results suggest that the success of these techniques depends heavily on the methods used for training dataset construction, but with proper consideration of representation bias, they can reduce the disparity seen among dialect subgroups on the hate speech detection task.
title Evaluating Simple Debiasing Techniques in RoBERTa-based Hate Speech Detection Models
topic Computation and Language
url https://arxiv.org/abs/2501.15430