Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tekanlou, Hadi Bayrami Asl, Bakhtiyarzadeh, Mahdi, Razmara, Jafar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914621569368064
author Tekanlou, Hadi Bayrami Asl
Bakhtiyarzadeh, Mahdi
Razmara, Jafar
author_facet Tekanlou, Hadi Bayrami Asl
Bakhtiyarzadeh, Mahdi
Razmara, Jafar
contents The spread of hate speech has become increasingly harmful in modern digital environments, particularly on social networking platforms. While recent advances have shown promising results in automatic hate speech detection, a key challenge remains: distinguishing genuine hate speech from reclaimed language. Accurate labeling is difficult due to the nuanced and context-dependent nature of reclaimed expressions. In this paper, we present a simple and interpretable approach for distinguishing hate speech from reclaimed language, developed for the MultiPride Shared Task. Our method generates dense semantic text embeddings and incorporates a label-noise filtering stage using Cleanlab with logistic regression, followed by a Multi-layer Perceptron (MLP) neural network for final classification. The system is designed to operate under limited computational resources while maintaining strong performance. We evaluate our approach using precision, recall, and F1-score, including macro-averaged metrics. Experimental results demonstrate robust performance despite extreme class imbalance in the dataset. Overall, the findings highlight the potential for further improvements through larger embedding models and more advanced preprocessing techniques while preserving interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01298
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?
Tekanlou, Hadi Bayrami Asl
Bakhtiyarzadeh, Mahdi
Razmara, Jafar
Computation and Language
The spread of hate speech has become increasingly harmful in modern digital environments, particularly on social networking platforms. While recent advances have shown promising results in automatic hate speech detection, a key challenge remains: distinguishing genuine hate speech from reclaimed language. Accurate labeling is difficult due to the nuanced and context-dependent nature of reclaimed expressions. In this paper, we present a simple and interpretable approach for distinguishing hate speech from reclaimed language, developed for the MultiPride Shared Task. Our method generates dense semantic text embeddings and incorporates a label-noise filtering stage using Cleanlab with logistic regression, followed by a Multi-layer Perceptron (MLP) neural network for final classification. The system is designed to operate under limited computational resources while maintaining strong performance. We evaluate our approach using precision, recall, and F1-score, including macro-averaged metrics. Experimental results demonstrate robust performance despite extreme class imbalance in the dataset. Overall, the findings highlight the potential for further improvements through larger embedding models and more advanced preprocessing techniques while preserving interpretability.
title Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?
topic Computation and Language
url https://arxiv.org/abs/2606.01298