BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917056534806528 |
|---|---|
| author | Mohsin, Ayesha Afroza Ahsan, Mashrur Maliyat, Nafisa Maria, Shanta Raiyan, Syed Rifat Mahmud, Hasan Hasan, Md Kamrul |
| author_facet | Mohsin, Ayesha Afroza Ahsan, Mashrur Maliyat, Nafisa Maria, Shanta Raiyan, Syed Rifat Mahmud, Hasan Hasan, Md Kamrul |
| contents | Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains underexplored due to limited resources. In this paper, we propose a novel pipeline for Bengali text detoxification that combines Pareto class-optimized large language models (LLMs) and Chain-of-Thought (CoT) prompting to generate detoxified sentences. To support this effort, we construct BanglaNirTox, an artificially generated parallel corpus of 68,041 toxic Bengali sentences with class-wise toxicity labels, reasonings, and detoxified paraphrases, using Pareto-optimized LLMs evaluated on random samples. The resulting BanglaNirTox dataset is used to fine-tune language models to produce better detoxified versions of Bengali sentences. Our findings show that Pareto-optimized LLMs with CoT prompting significantly enhance the quality and consistency of Bengali text detoxification. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_01512 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification Mohsin, Ayesha Afroza Ahsan, Mashrur Maliyat, Nafisa Maria, Shanta Raiyan, Syed Rifat Mahmud, Hasan Hasan, Md Kamrul Computation and Language Artificial Intelligence Toxic language in Bengali remains prevalent, especially in online environments, with few effective precautions against it. Although text detoxification has seen progress in high-resource languages, Bengali remains underexplored due to limited resources. In this paper, we propose a novel pipeline for Bengali text detoxification that combines Pareto class-optimized large language models (LLMs) and Chain-of-Thought (CoT) prompting to generate detoxified sentences. To support this effort, we construct BanglaNirTox, an artificially generated parallel corpus of 68,041 toxic Bengali sentences with class-wise toxicity labels, reasonings, and detoxified paraphrases, using Pareto-optimized LLMs evaluated on random samples. The resulting BanglaNirTox dataset is used to fine-tune language models to produce better detoxified versions of Bengali sentences. Our findings show that Pareto-optimized LLMs with CoT prompting significantly enhance the quality and consistency of Bengali text detoxification. |
| title | BanglaNirTox: A Large-scale Parallel Corpus for Explainable AI in Bengali Text Detoxification |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2511.01512 |