Decoding Islamophobic Discourse: Using LLMs to Identify Tropes and Semi-Coded Hate Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mustafa, Raza Ul, Dupart, Roi, Smith, Gabrielle, Ashraf, Noman, Japkowicz, Nathalie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915996504162304
author Mustafa, Raza Ul
Dupart, Roi
Smith, Gabrielle
Ashraf, Noman
Japkowicz, Nathalie
author_facet Mustafa, Raza Ul
Dupart, Roi
Smith, Gabrielle
Ashraf, Noman
Japkowicz, Nathalie
contents In recent years, Islamophobia has gained significant traction across Western societies, fueled by the rise of digital communication networks. This paper performs a large-scale analysis of specialized, semi-coded Islamophobic terms such as (muzrat, pislam, mudslime, mohammedan, muzzies) floated on extremist social platforms, i.e., 4Chan, Gab, Telegram, etc. Many of these terms appear lexically neutral or ambiguous outside of specific contexts, making them difficult for both human moderators and automated systems to reliably identify as hate speech. First, we use Large Language Models (LLMs) to show their ability to understand these terms. Second, Google Perspective API suggests that Islamophobic posts tend to receive higher toxicity scores than other categories of hate speech like Antisemitism. Finally, we use BERT topic modeling approach to extract different topics and Islamophobic discourse on these social platforms. Our findings indicate that LLMs understand these Out-Of-Vocabulary (OOV) slurs; however, further improvements in moderation strategies and algorithmic detection are necessary to address such discourse effectively. Our topic modeling also indicates that Islamophobic text is found across various political, conspiratorial, and far-right movements and is particularly directed against Muslim immigrants. Taken altogether, we performed one of the first studies on Islamophobic semi-coded terms and shed a global light on Islamophobia.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18273
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Decoding Islamophobic Discourse: Using LLMs to Identify Tropes and Semi-Coded Hate Speech
Mustafa, Raza Ul
Dupart, Roi
Smith, Gabrielle
Ashraf, Noman
Japkowicz, Nathalie
Machine Learning
In recent years, Islamophobia has gained significant traction across Western societies, fueled by the rise of digital communication networks. This paper performs a large-scale analysis of specialized, semi-coded Islamophobic terms such as (muzrat, pislam, mudslime, mohammedan, muzzies) floated on extremist social platforms, i.e., 4Chan, Gab, Telegram, etc. Many of these terms appear lexically neutral or ambiguous outside of specific contexts, making them difficult for both human moderators and automated systems to reliably identify as hate speech. First, we use Large Language Models (LLMs) to show their ability to understand these terms. Second, Google Perspective API suggests that Islamophobic posts tend to receive higher toxicity scores than other categories of hate speech like Antisemitism. Finally, we use BERT topic modeling approach to extract different topics and Islamophobic discourse on these social platforms. Our findings indicate that LLMs understand these Out-Of-Vocabulary (OOV) slurs; however, further improvements in moderation strategies and algorithmic detection are necessary to address such discourse effectively. Our topic modeling also indicates that Islamophobic text is found across various political, conspiratorial, and far-right movements and is particularly directed against Muslim immigrants. Taken altogether, we performed one of the first studies on Islamophobic semi-coded terms and shed a global light on Islamophobia.
title Decoding Islamophobic Discourse: Using LLMs to Identify Tropes and Semi-Coded Hate Speech
topic Machine Learning
url https://arxiv.org/abs/2503.18273