BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Haque, Farah Binta, Yasin, Md, Saha, Shishir, Rafi, Md Shoaib Akhter, Sadeque, Farig
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918198193946624
author Haque, Farah Binta
Yasin, Md
Saha, Shishir
Rafi, Md Shoaib Akhter
Sadeque, Farig
author_facet Haque, Farah Binta
Yasin, Md
Saha, Shishir
Rafi, Md Shoaib Akhter
Sadeque, Farig
contents Despite the growing progress in Natural Language Inference (NLI) research, resources for the Bengali language remain extremely limited. Existing Bengali NLI datasets exhibit several inconsistencies, including annotation errors, ambiguous sentence pairs, and inadequate linguistic diversity, which hinder effective model training and evaluation. To address these limitations, we introduce BNLI, a refined and linguistically curated Bengali NLI dataset designed to support robust language understanding and inference modeling. The dataset was constructed through a rigorous annotation pipeline emphasizing semantic clarity and balance across entailment, contradiction, and neutrality classes. We benchmarked BNLI using a suite of state-of-the-art transformer-based architectures, including multilingual and Bengali-specific models, to assess their ability to capture complex semantic relations in Bengali text. The experimental findings highlight the improved reliability and interpretability achieved with BNLI, establishing it as a strong foundation for advancing research in Bengali and other low-resource language inference tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2511_08813
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference
Haque, Farah Binta
Yasin, Md
Saha, Shishir
Rafi, Md Shoaib Akhter
Sadeque, Farig
Computation and Language
Despite the growing progress in Natural Language Inference (NLI) research, resources for the Bengali language remain extremely limited. Existing Bengali NLI datasets exhibit several inconsistencies, including annotation errors, ambiguous sentence pairs, and inadequate linguistic diversity, which hinder effective model training and evaluation. To address these limitations, we introduce BNLI, a refined and linguistically curated Bengali NLI dataset designed to support robust language understanding and inference modeling. The dataset was constructed through a rigorous annotation pipeline emphasizing semantic clarity and balance across entailment, contradiction, and neutrality classes. We benchmarked BNLI using a suite of state-of-the-art transformer-based architectures, including multilingual and Bengali-specific models, to assess their ability to capture complex semantic relations in Bengali text. The experimental findings highlight the improved reliability and interpretability achieved with BNLI, establishing it as a strong foundation for advancing research in Bengali and other low-resource language inference tasks.
title BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference
topic Computation and Language
url https://arxiv.org/abs/2511.08813