BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fayaz, Azizul Hakim, Uddin, MD. Shorif, Bhuiyan, Rayhan Uddin, Sultana, Zakia, Islam, Md. Samiul, Paul, Bidyarthi, Muhammad, Tashreef, Manzoor, Shahriar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911069572694016
author Fayaz, Azizul Hakim
Uddin, MD. Shorif
Bhuiyan, Rayhan Uddin
Sultana, Zakia
Islam, Md. Samiul
Paul, Bidyarthi
Muhammad, Tashreef
Manzoor, Shahriar
author_facet Fayaz, Azizul Hakim
Uddin, MD. Shorif
Bhuiyan, Rayhan Uddin
Sultana, Zakia
Islam, Md. Samiul
Paul, Bidyarthi
Muhammad, Tashreef
Manzoor, Shahriar
contents Hate speech on digital platforms has become a growing concern globally, especially in linguistically diverse countries like Bangladesh, where regional dialects play a major role in everyday communication. Despite progress in hate speech detection for standard Bangla, Existing datasets and systems fail to address the informal and culturally rich expressions found in dialects such as Barishal, Noakhali, and Chittagong. This oversight results in limited detection capability and biased moderation, leaving large sections of harmful content unaccounted for. To address this gap, this study introduces BIDWESH, the first multi-dialectal Bangla hate speech dataset, constructed by translating and annotating 9,183 instances from the BD-SHS corpus into three major regional dialects. Each entry was manually verified and labeled for hate presence, type (slander, gender, religion, call to violence), and target group (individual, male, female, group), ensuring linguistic and contextual accuracy. The resulting dataset provides a linguistically rich, balanced, and inclusive resource for advancing hate speech detection in Bangla. BIDWESH lays the groundwork for the development of dialect-sensitive NLP tools and contributes significantly to equitable and context-aware content moderation in low-resource language settings.
format Preprint
id arxiv_https___arxiv_org_abs_2507_16183
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset
Fayaz, Azizul Hakim
Uddin, MD. Shorif
Bhuiyan, Rayhan Uddin
Sultana, Zakia
Islam, Md. Samiul
Paul, Bidyarthi
Muhammad, Tashreef
Manzoor, Shahriar
Computation and Language
Hate speech on digital platforms has become a growing concern globally, especially in linguistically diverse countries like Bangladesh, where regional dialects play a major role in everyday communication. Despite progress in hate speech detection for standard Bangla, Existing datasets and systems fail to address the informal and culturally rich expressions found in dialects such as Barishal, Noakhali, and Chittagong. This oversight results in limited detection capability and biased moderation, leaving large sections of harmful content unaccounted for. To address this gap, this study introduces BIDWESH, the first multi-dialectal Bangla hate speech dataset, constructed by translating and annotating 9,183 instances from the BD-SHS corpus into three major regional dialects. Each entry was manually verified and labeled for hate presence, type (slander, gender, religion, call to violence), and target group (individual, male, female, group), ensuring linguistic and contextual accuracy. The resulting dataset provides a linguistically rich, balanced, and inclusive resource for advancing hate speech detection in Bangla. BIDWESH lays the groundwork for the development of dialect-sensitive NLP tools and contributes significantly to equitable and context-aware content moderation in low-resource language settings.
title BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset
topic Computation and Language
url https://arxiv.org/abs/2507.16183