BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hasan, Md. Najib, Rain, Mst. Jannatun Ferdous, Mohammed, Fyad, Siddique, Nazmul
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911462377652224
author Hasan, Md. Najib
Rain, Mst. Jannatun Ferdous
Mohammed, Fyad
Siddique, Nazmul
author_facet Hasan, Md. Najib
Rain, Mst. Jannatun Ferdous
Mohammed, Fyad
Siddique, Nazmul
contents IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensive and difficult to scale, while using large language models (LLMs) as automated annotators introduces concerns about label reliability, bias, and evaluation validity. This work presents a Bangla IR dataset constructed using a BETA-labeling framework involving multiple LLM annotators from diverse model families. The framework incorporates contextual alignment, consistency checks, and majority agreement, followed by human evaluation to verify label quality. Beyond dataset creation, we examine whether IR datasets from other low-resource languages can be effectively reused through one-hop machine translation. Using LLM-based translation across multiple language pairs, we experimented on meaning preservation and task validity between source and translated datasets. Our experiment reveal substantial variation across languages, reflecting language-dependent biases and inconsistent semantic preservation that directly affect the reliability of cross-lingual dataset reuse. Overall, this study highlights both the potential and limitations of LLM-assisted dataset creation for low-resource IR. It provides empirical evidence of the risks associated with cross-lingual dataset reuse and offers practical guidance for constructing more reliable benchmarks and evaluation pipelines in low-resource language settings.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14488
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR
Hasan, Md. Najib
Rain, Mst. Jannatun Ferdous
Mohammed, Fyad
Siddique, Nazmul
Computation and Language
Artificial Intelligence
IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensive and difficult to scale, while using large language models (LLMs) as automated annotators introduces concerns about label reliability, bias, and evaluation validity. This work presents a Bangla IR dataset constructed using a BETA-labeling framework involving multiple LLM annotators from diverse model families. The framework incorporates contextual alignment, consistency checks, and majority agreement, followed by human evaluation to verify label quality. Beyond dataset creation, we examine whether IR datasets from other low-resource languages can be effectively reused through one-hop machine translation. Using LLM-based translation across multiple language pairs, we experimented on meaning preservation and task validity between source and translated datasets. Our experiment reveal substantial variation across languages, reflecting language-dependent biases and inconsistent semantic preservation that directly affect the reliability of cross-lingual dataset reuse. Overall, this study highlights both the potential and limitations of LLM-assisted dataset creation for low-resource IR. It provides empirical evidence of the risks associated with cross-lingual dataset reuse and offers practical guidance for constructing more reliable benchmarks and evaluation pipelines in low-resource language settings.
title BETA-Labeling for Multilingual Dataset Construction in Low-Resource IR
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2602.14488