BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Mohsin, Ayesha Afroza, Ahsan, Mashrur, Maliyat, Nafisa, Maria, Shanta, Raiyan, Syed Rifat, Mahmud, Hasan, Hasan, Md Kamrul
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917056534806528
author Mohsin, Ayesha Afroza
Ahsan, Mashrur
Maliyat, Nafisa
Maria, Shanta
Raiyan, Syed Rifat
Mahmud, Hasan
Hasan, Md Kamrul
author_facet Mohsin, Ayesha Afroza
Ahsan, Mashrur
Maliyat, Nafisa
Maria, Shanta
Raiyan, Syed Rifat
Mahmud, Hasan
Hasan, Md Kamrul
contents Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains underexplored due to limited resources. In this paper, we propose a novel pipeline for Bengali text detoxification that combines Pareto class-optimized large language models (LLMs) and Chain-of-Thought (CoT) prompting to generate detoxified sentences. To support this effort, we construct BanglaNirTox, an artificially generated parallel corpus of 68,041 toxic Bengali sentences with class-wise toxicity labels, reasonings, and detoxified paraphrases, using Pareto-optimized LLMs evaluated on random samples. The resulting BanglaNirTox dataset is used to fine-tune language models to produce better detoxified versions of Bengali sentences. Our findings show that Pareto-optimized LLMs with CoT prompting significantly enhance the quality and consistency of Bengali text detoxification.
format Preprint
id arxiv_https___arxiv_org_abs_2511_01512
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification
Mohsin, Ayesha Afroza
Ahsan, Mashrur
Maliyat, Nafisa
Maria, Shanta
Raiyan, Syed Rifat
Mahmud, Hasan
Hasan, Md Kamrul
Computation and Language
Artificial Intelligence
Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains underexplored due to limited resources. In this paper, we propose a novel pipeline for Bengali text detoxification that combines Pareto class-optimized large language models (LLMs) and Chain-of-Thought (CoT) prompting to generate detoxified sentences. To support this effort, we construct BanglaNirTox, an artificially generated parallel corpus of 68,041 toxic Bengali sentences with class-wise toxicity labels, reasonings, and detoxified paraphrases, using Pareto-optimized LLMs evaluated on random samples. The resulting BanglaNirTox dataset is used to fine-tune language models to produce better detoxified versions of Bengali sentences. Our findings show that Pareto-optimized LLMs with CoT prompting significantly enhance the quality and consistency of Bengali text detoxification.
title BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.01512