Who Speaks Matters: Analysing the Influence of the Speaker's Ethnicity on Hate Classification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Malik, Ananya, Sharma, Kartik, Bhatt, Shaily, Ng, Lynnette Hui Xian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915547207172096
author Malik, Ananya
Sharma, Kartik
Bhatt, Shaily
Ng, Lynnette Hui Xian
author_facet Malik, Ananya
Sharma, Kartik
Bhatt, Shaily
Ng, Lynnette Hui Xian
contents Large Language Models (LLMs) offer a lucrative promise for scalable content moderation, including hate speech detection. However, they are also known to be brittle and biased against marginalised communities and dialects. This requires their applications to high-stakes tasks like hate speech detection to be critically scrutinized. In this work, we investigate the robustness of hate speech classification using LLMs particularly when explicit and implicit markers of the speaker's ethnicity are injected into the input. For explicit markers, we inject a phrase that mentions the speaker's linguistic identity. For the implicit markers, we inject dialectal features. By analysing how frequently model outputs flip in the presence of these markers, we reveal varying degrees of brittleness across 3 LLMs and 1 LM and 5 linguistic identities. We find that the presence of implicit dialect markers in inputs causes model outputs to flip more than the presence of explicit markers. Further, the percentage of flips varies across ethnicities. Finally, we find that larger models are more robust. Our findings indicate the need for exercising caution in deploying LLMs for high-stakes tasks like hate speech detection.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20490
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Who Speaks Matters: Analysing the Influence of the Speaker's Ethnicity on Hate Classification
Malik, Ananya
Sharma, Kartik
Bhatt, Shaily
Ng, Lynnette Hui Xian
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) offer a lucrative promise for scalable content moderation, including hate speech detection. However, they are also known to be brittle and biased against marginalised communities and dialects. This requires their applications to high-stakes tasks like hate speech detection to be critically scrutinized. In this work, we investigate the robustness of hate speech classification using LLMs particularly when explicit and implicit markers of the speaker's ethnicity are injected into the input. For explicit markers, we inject a phrase that mentions the speaker's linguistic identity. For the implicit markers, we inject dialectal features. By analysing how frequently model outputs flip in the presence of these markers, we reveal varying degrees of brittleness across 3 LLMs and 1 LM and 5 linguistic identities. We find that the presence of implicit dialect markers in inputs causes model outputs to flip more than the presence of explicit markers. Further, the percentage of flips varies across ethnicities. Finally, we find that larger models are more robust. Our findings indicate the need for exercising caution in deploying LLMs for high-stakes tasks like hate speech detection.
title Who Speaks Matters: Analysing the Influence of the Speaker's Ethnicity on Hate Classification
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.20490