Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hosain, Md. Tanzib, Gupta, Rajan Das, Morol, Md. Kishor
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908384148586496
author Hosain, Md. Tanzib
Gupta, Rajan Das
Morol, Md. Kishor
author_facet Hosain, Md. Tanzib
Gupta, Rajan Das
Morol, Md. Kishor
contents In this work, we provide DZEN, a dataset of parallel Dzongkha and English test questions for Bhutanese middle and high school students. The over 5K questions in our collection span a variety of scientific topics and include factual, application, and reasoning-based questions. We use our parallel dataset to test a number of Large Language Models (LLMs) and find a significant performance difference between the models in English and Dzongkha. We also look at different prompting strategies and discover that Chain-of-Thought (CoT) prompting works well for reasoning questions but less well for factual ones. We also find that adding English translations enhances the precision of Dzongkha question responses. Our results point to exciting avenues for further study to improve LLM performance in Dzongkha and, more generally, in low-resource languages. We release the dataset at: https://github.com/kraritt/llm_dzongkha_evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models
Hosain, Md. Tanzib
Gupta, Rajan Das
Morol, Md. Kishor
Computation and Language
In this work, we provide DZEN, a dataset of parallel Dzongkha and English test questions for Bhutanese middle and high school students. The over 5K questions in our collection span a variety of scientific topics and include factual, application, and reasoning-based questions. We use our parallel dataset to test a number of Large Language Models (LLMs) and find a significant performance difference between the models in English and Dzongkha. We also look at different prompting strategies and discover that Chain-of-Thought (CoT) prompting works well for reasoning questions but less well for factual ones. We also find that adding English translations enhances the precision of Dzongkha question responses. Our results point to exciting avenues for further study to improve LLM performance in Dzongkha and, more generally, in low-resource languages. We release the dataset at: https://github.com/kraritt/llm_dzongkha_evaluation.
title Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models
topic Computation and Language
url https://arxiv.org/abs/2505.18638