DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lucas, Jason, Murtagh, Matt, Al-Lawati, Ali, Uchendu, Uchendu, Uchendu, Adaku, Lee, Dongwon
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918430660100096
author Lucas, Jason
Murtagh, Matt
Al-Lawati, Ali
Uchendu, Uchendu
Uchendu, Adaku
Lee, Dongwon
author_facet Lucas, Jason
Murtagh, Matt
Al-Lawati, Ali
Uchendu, Uchendu
Uchendu, Adaku
Lee, Dongwon
contents Harmful content detectors-particularly disinformation classifiers-are predominantly developed and evaluated on Standard American English (SAE), leaving their robustness to dialectal variation unexplored. We present DIA-HARM, the first benchmark for evaluating disinformation detection robustness across 50 English dialects spanning U.S., British, African, Caribbean, and Asia-Pacific varieties. Using Multi-VALUE's linguistically grounded transformations, we introduce D3 (Dialectal Disinformation Detection), a corpus of 195K samples derived from established disinformation benchmarks. Our evaluation of 16 detection models reveals systematic vulnerabilities: human-written dialectal content degrades detection by 1.4-3.6% F1, while AI-generated content remains stable. Fine-tuned transformers substantially outperform zero-shot LLMs (96.6% vs. 78.3% best-case F1), with some models exhibiting catastrophic failures exceeding 33% degradation on mixed content. Cross-dialectal transfer analysis across 2,450 dialect pairs shows that multilingual models (mDeBERTa: 97.2% average F1) generalize effectively, while monolingual models like RoBERTa and XLM-RoBERTa fail on dialectal inputs. These findings demonstrate that current disinformation detectors may systematically disadvantage hundreds of millions of non-SAE speakers worldwide. We release the DIA-HARM framework, D3 corpus, and evaluation tools: https://github.com/jsl5710/dia-harm
format Preprint
id arxiv_https___arxiv_org_abs_2604_05318
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects
Lucas, Jason
Murtagh, Matt
Al-Lawati, Ali
Uchendu, Uchendu
Uchendu, Adaku
Lee, Dongwon
Computation and Language
Harmful content detectors-particularly disinformation classifiers-are predominantly developed and evaluated on Standard American English (SAE), leaving their robustness to dialectal variation unexplored. We present DIA-HARM, the first benchmark for evaluating disinformation detection robustness across 50 English dialects spanning U.S., British, African, Caribbean, and Asia-Pacific varieties. Using Multi-VALUE's linguistically grounded transformations, we introduce D3 (Dialectal Disinformation Detection), a corpus of 195K samples derived from established disinformation benchmarks. Our evaluation of 16 detection models reveals systematic vulnerabilities: human-written dialectal content degrades detection by 1.4-3.6% F1, while AI-generated content remains stable. Fine-tuned transformers substantially outperform zero-shot LLMs (96.6% vs. 78.3% best-case F1), with some models exhibiting catastrophic failures exceeding 33% degradation on mixed content. Cross-dialectal transfer analysis across 2,450 dialect pairs shows that multilingual models (mDeBERTa: 97.2% average F1) generalize effectively, while monolingual models like RoBERTa and XLM-RoBERTa fail on dialectal inputs. These findings demonstrate that current disinformation detectors may systematically disadvantage hundreds of millions of non-SAE speakers worldwide. We release the DIA-HARM framework, D3 corpus, and evaluation tools: https://github.com/jsl5710/dia-harm
title DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects
topic Computation and Language
url https://arxiv.org/abs/2604.05318