DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects
Fuente:
arXiv
Saved in:
| Main Authors: | Lucas, Jason, Murtagh, Matt, Al-Lawati, Ali, Uchendu, Uchendu, Uchendu, Adaku, Lee, Dongwon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages
by: Lucas, Jason, et al.
Published: (2026)
by: Lucas, Jason, et al.
Published: (2026)
GPT-who: An Information Density-based Machine-Generated Text Detector
by: Venkatraman, Saranya, et al.
Published: (2023)
by: Venkatraman, Saranya, et al.
Published: (2023)
TOPFORMER: Topology-Aware Authorship Attribution of Deepfake Texts with Diverse Writing Styles
by: Uchendu, Adaku, et al.
Published: (2023)
by: Uchendu, Adaku, et al.
Published: (2023)
Topological Data Analysis Applications in Natural Language Processing: A Survey
by: Uchendu, Adaku, et al.
Published: (2024)
by: Uchendu, Adaku, et al.
Published: (2024)
Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
by: Guo, William, et al.
Published: (2025)
by: Guo, William, et al.
Published: (2025)
PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection
by: Lee, Jooyoung, et al.
Published: (2024)
by: Lee, Jooyoung, et al.
Published: (2024)
Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
Beemo: Benchmark of Expert-edited Machine-generated Outputs
by: Artemova, Ekaterina, et al.
Published: (2024)
by: Artemova, Ekaterina, et al.
Published: (2024)
Authorship Obfuscation in Multilingual Machine-Generated Text Detection
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
by: Macko, Dominik, et al.
Published: (2023)
by: Macko, Dominik, et al.
Published: (2023)
Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification
by: Alperin, Kenneth, et al.
Published: (2025)
by: Alperin, Kenneth, et al.
Published: (2025)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
by: Dorn, Rebecca, et al.
Published: (2024)
by: Dorn, Rebecca, et al.
Published: (2024)
Exploiting Dialect Identification in Automatic Dialectal Text Normalization
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
Semantic Captioning: Benchmark Dataset and Graph-Aware Few-Shot In-Context Learning for SQL2Text
by: Al-Lawati, Ali, et al.
Published: (2025)
by: Al-Lawati, Ali, et al.
Published: (2025)
Computational Linguistics Meets Libyan Dialect: A Study on Dialect Identification
by: Essgaer, Mansour, et al.
Published: (2025)
by: Essgaer, Mansour, et al.
Published: (2025)
DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English
by: Oh, Jio, et al.
Published: (2026)
by: Oh, Jio, et al.
Published: (2026)
Algerian Dialect
by: Benmounah, Zakaria, et al.
Published: (2025)
by: Benmounah, Zakaria, et al.
Published: (2025)
A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
The Hrunting of AI: Where and How to Improve English Dialectal Fairness
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
Improving Dialectal Slot and Intent Detection with Auxiliary Tasks: A Multi-Dialectal Bavarian Case Study
by: Krückl, Xaver Maria, et al.
Published: (2025)
by: Krückl, Xaver Maria, et al.
Published: (2025)
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties
by: Faisal, Fahim, et al.
Published: (2024)
by: Faisal, Fahim, et al.
Published: (2024)
Dolphin-CN-Dialect: Where Chinese Dialects Matter
by: Meng, Yangyang, et al.
Published: (2026)
by: Meng, Yangyang, et al.
Published: (2026)
A Catalog of Basque Dialectal Resources: Online Collections and Standard-to-Dialectal Adaptations
by: Bengoetxea, Jaione, et al.
Published: (2026)
by: Bengoetxea, Jaione, et al.
Published: (2026)
ArabicDialectHub: A Cross-Dialectal Arabic Learning Resource and Platform
by: Lahlou, Salem
Published: (2026)
by: Lahlou, Salem
Published: (2026)
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models
by: Bafna, Niyati, et al.
Published: (2025)
by: Bafna, Niyati, et al.
Published: (2025)
JEEM: Vision-Language Understanding in Four Arabic Dialects
by: Kadaoui, Karima, et al.
Published: (2025)
by: Kadaoui, Karima, et al.
Published: (2025)
A Novel Dialect-Aware Framework for the Classification of Arabic Dialects and Emotions
by: Alsadhan, Nasser A
Published: (2025)
by: Alsadhan, Nasser A
Published: (2025)
Extracting Lexical Features from Dialects via Interpretable Dialect Classifiers
by: Xie, Roy, et al.
Published: (2024)
by: Xie, Roy, et al.
Published: (2024)
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
by: Zhou, Yu, et al.
Published: (2025)
by: Zhou, Yu, et al.
Published: (2025)
Phonotactic Complexity across Dialects
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
Graph-based Molecular In-context Learning Grounded on Morgan Fingerprints
by: Al-Lawati, Ali, et al.
Published: (2025)
by: Al-Lawati, Ali, et al.
Published: (2025)
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
by: Altakrori, Malik H., et al.
Published: (2025)
by: Altakrori, Malik H., et al.
Published: (2025)
Harnessing Test-time Adaptation for NLU tasks Involving Dialects of English
by: Nguyen, Duke, et al.
Published: (2025)
by: Nguyen, Duke, et al.
Published: (2025)
ArFake: A Multi-Dialect Benchmark and Baselines for Arabic Spoof-Speech Detection
by: Maged, Mohamed, et al.
Published: (2025)
by: Maged, Mohamed, et al.
Published: (2025)
A Survey of Large Language Models for Arabic Language and its Dialects
by: Mashaabi, Malak, et al.
Published: (2024)
by: Mashaabi, Malak, et al.
Published: (2024)
Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects
by: Sayeedi, Nurul Labib, et al.
Published: (2026)
by: Sayeedi, Nurul Labib, et al.
Published: (2026)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
Mitigating Harmful Erraticism in LLMs Through Dialectical Behavior Therapy Based De-Escalation Strategies
by: Rangarajan, Pooja, et al.
Published: (2025)
by: Rangarajan, Pooja, et al.
Published: (2025)
Saudi-Dialect-ALLaM: LoRA Fine-Tuning for Dialectal Arabic Generation
by: Barmandah, Hassan
Published: (2025)
by: Barmandah, Hassan
Published: (2025)
Similar Items
-
BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages
by: Lucas, Jason, et al.
Published: (2026) -
GPT-who: An Information Density-based Machine-Generated Text Detector
by: Venkatraman, Saranya, et al.
Published: (2023) -
TOPFORMER: Topology-Aware Authorship Attribution of Deepfake Texts with Diverse Writing Styles
by: Uchendu, Adaku, et al.
Published: (2023) -
Topological Data Analysis Applications in Natural Language Processing: A Survey
by: Uchendu, Adaku, et al.
Published: (2024) -
Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
by: Guo, William, et al.
Published: (2025)